Market Analysis8 August 2026

ShipItClean and the audit-capacity gap

AI is writing code faster than anyone can review it, at a measurably higher defect rate, and the consequences have started appearing in the public vulnerability record. This is an analysis of the market that gap creates, of where ShipItClean is actually positioned within it, and of the four numbers still missing before the case can be underwritten.

Authors: Larry Arnold | Apollo Raines

Scope. This document sizes and structures a market. It deliberately contains no company valuation, no raise terms and no price per share. Those live in the separate valuation work and are best assessed against comparables rather than derived from anything here.

Provenance flags. Every quantitative claim is marked, because the sources vary in quality and several disagree.

Cited Sourced to a named public source, listed at the end.

Range Sources materially disagree. The span is shown rather than a midpoint; treat the spread as the finding.

Internal From our own product and valuation work, not independently verified here.

The thesis in short

The demand driver is real, measurable and accelerating. AI now authors a substantial and growing share of production code; that code carries a materially higher defect rate than human-written code; and AI-attributable CVEs have begun compounding month over month. Review demand scales with volume multiplied by defect rate. Human audit capacity is flat.

What makes ShipItClean interesting is not that it addresses this gap, because a dozen funded companies do. It is that it addresses the part of the gap the others structurally cannot reach, and does so at a cost base nobody else in the category has.

Central argument

Every competitor reviews the code that changed. ShipItClean audits the code that exists. That is a different unit of work, sold on different economics, and it is only possible because Atlas removes the context ceiling that bounds everyone else. It has been demonstrated at 44 million tokens of Firefox source for roughly $15 of electricity, against a competitor quote near $50,000. The full report is public.

A note on brand. The adversarial audit product was originally marketed as Hostile Review but has been consolidated under the ShipItClean brand; earlier references to that name may still appear in external coverage. Internal

Which market this actually is

The most common analytical error with ShipItClean, and the one made in the first draft of this document, is to size it as an AI code review company. It is worth being precise about why that is wrong, because the mistake changes the answer by an order of magnitude and points the go-to-market at the wrong buyer.

Two different units of work
 AI code reviewAdversarial audit
UnitThe pull request diffThe whole codebase
CadenceContinuous, every PRPre-release, quarterly, on acquisition or incident
PricingPer developer seat, $24-48/month CitedPer scan; team size irrelevant Internal
BuyerEngineering lead, bottom-upSecurity, compliance, or the board after an incident
DisplacesReviewer attentionManual audit and penetration-test engagements
PlayersCodeRabbit, Greptile, Qodo, Copilot reviewAudit firms, pentest retainers, internal AppSec

ShipItClean's own comparison pages hold this line consistently: pay per scan, no seats, a solo developer and a two-hundred-person team pay the same rate. That is not a pricing preference. It is a statement that the product is not sold by the seat because it does not do per-seat work.

A second commercial tier is live alongside the per-scan product: a conversational interface where an engineer interrogates a model that holds the entire codebase in Atlas memory, sold on monthly or annual contracts. Internal That tier reprices ShipItClean from a transactional scan into recurring revenue, and changes the buyer conversation from "pay per audit" to "retain a full-context security partner."

Consequence for sizing

The AI code review market is the category ShipItClean is compared to and recruits users from. It is not the category it monetises. Sizing on developer seats produces a confident number for the wrong business.

Sizing

The reference market: adjacent, not target

Pure-play AI code review vendors collectively run at roughly $420M ARR in 2026, up from about $180M in 2025. Cited Set against approximately 28.7 million professional developers at prevailing seat prices, that implies the seat market is running at roughly 5% penetration. Range

This matters as context in two ways and no others: it establishes that the adjacent category is early rather than saturated, and it is where the free tier recruits from. It is not ShipItClean's addressable market.

The target market

Analyst-tracked markets ShipItClean sits across
MarketCurrentForecastGrowthFlag
Application security testing (SAST, DAST, SCA)$1.83B (2025)$7.60B by 203126.7% CAGRCited
Application security (broad)~$13.6B$28.1B by 2031n/aInternal
DevSecOps~$9-11Bn/an/aInternal
AI code tools (broad)$9.35-9.8B (2026)$26.0B by 203026-27% CAGRCited
SAST alone$2.8-4.0B$6.3-15.5B10.8-24.1% CAGRRange

Framed conservatively as application security testing plus AI code review, TAM is roughly $10-15B today, scaling past $40B by the early 2030s. Internal The serviceable slice is enterprises and public-sector bodies with large, sensitive or regulated codebases: precisely where whole-system context is worth paying for, and where per-seat tools have nothing to offer.

On the SAST spread. Published 2026 SAST estimates run $2.8B to $4.0B, with growth between 10.8% and 24.1%. The high case more than doubles the low over the same horizon. That is a market-definition disagreement rather than a forecasting one, and any figure quoted from this row without its spread is being quoted dishonestly.

Who the SAM actually contains

"Large, sensitive or regulated codebases" is the stated serviceable market. Three archetypes make it concrete:

  • Industrial and energy majors (Chevron-scale operators): safety-critical codebases, SCADA, refinery control and trading systems where a single missed defect is a catastrophic-loss event. Whole-system reasoning is not a convenience; it is the only way to trace trust boundaries across those systems. Internal
  • Hyperscalers (Microsoft-scale platforms): ship their own AI coding tools yet still need independent, full-context review of codebases so large that only an uncapped context substrate can hold them.
  • Federal government and regulated public sector: procurement mandates, FedRAMP-style controls and secure-software-development requirements create the strongest tailwind for a mandated, benchmarkable scanning standard. The sovereignty story of an owned model, where no code or prompts leave to a third-party API, is uniquely suited to classified and controlled environments. Internal

Bottom-up: the displacement anchor

Because the product is priced per scan, the bottom-up build is a displacement calculation rather than a seat count. The anchor is unusually clean: for the Firefox audit, a competitor quoted approximately $50,000 for work ShipItClean performed for about $15 in electricity. Internal

Three inputs convert that into a defensible obtainable-market figure, and all three are currently missing:

  1. Enterprise price per scan. Published credit pricing covers the self-serve tier only. Enterprise scan pricing is not established.
  2. Audit cadence per customer per year. Pre-release, quarterly and incident-driven scans imply very different revenue per account.
  3. Count of organisations with codebases that warrant it. The stated SAM, meaning large, sensitive or regulated codebases, has no headcount attached to it yet.

Those three numbers multiplied together are the obtainable market. Presenting a figure before they exist would be decoration, so none is offered here.

Why now: three curves

The timing argument rests on three independent series that have moved together over the past twelve months.

One: volume

The largest empirical study available, covering 4.2 million developers from November 2025 to February 2026, puts AI-authored production code at 26.9%. Copilot generates 46% of code in repositories where it is installed, against 27% in 2022; the tool has crossed 20 million users, is deployed across more than 90% of the Fortune 100, and grew users roughly 400% year over year. Cited CloudBees, on a broader definition, reports AI generating or assisting 61% of the average enterprise codebase. Google reported more than 30% of new code AI-generated as early as Q1 2025. Range

These figures are not in conflict. The spread is definitional: the low number measures code authored by AI, the high number measures code AI generated or assisted with. Quoting 61% as authorship, which is common in vendor material, is wrong. The defensible statement is that AI authorship of production code is above a quarter and rising.

Two: defect rate

Veracode, testing more than 100 models across 80 tasks, found 45% of AI-generated samples introduced an OWASP Top 10 vulnerability, with Java worst at 72%. Independent estimates of the multiple against human-written code range from 1.88x to 2.74x. Security pass rates sit near 55% despite syntactic correctness above 95%. GitClear found code churn roughly doubled from a 3.3% pre-AI baseline to 7.1%, with duplicated blocks up eightfold. A Stanford study found developers using AI assistants wrote less secure code while being more confident it was secure; an NYU study found roughly 40% of Copilot completions were vulnerable across MITRE Top-25 CWEs. Range

Three: consequence

Publicly disclosed CVEs directly attributable to AI-generated code went from 6 in January 2026, to 15 in February, to 35 in March. Cited This is the fastest-moving series in the dataset and the one that converts audit from discretionary to mandatory, and mandatory work is bought out of security budgets.

The bill is already large before any of this. CISQ puts the annual cost of poor software quality in the US at $2.41 trillion, and IBM puts the average US breach at $10.22M, with organisations running "shadow AI" incurring roughly $670,000 of additional breach cost each. Enterprises separately spend around $28,000 per developer per year on security-related engineering work. Cited

The Stanford finding is the sharpest one

Developers using AI assistants wrote less secure code and were more confident in it. Confidence moving in the opposite direction to quality is precisely the condition that self-review cannot fix, because the reviewer shares the author's blind spots. That is the argument for an independent adversary, and it is not a claim about ShipItClean's product quality, which makes it far harder to dismiss.

External validation: Mozilla and Mythos

In April 2026, Mozilla turned Anthropic's frontier model, Claude Mythos, loose on Firefox through a bespoke agentic harness bolted onto its existing fuzzing infrastructure. Firefox 150 shipped fixes for 271 AI-discovered vulnerabilities: 180 sec-high, 80 sec-moderate, 11 sec-low, some of them decades old, with near-zero false positives. Cited Firefox is one of the most-audited codebases on earth. If a hardened flagship is that full of AI-discoverable defects, the vastly larger universe of AI-written 2025-26 application code is the real, unaudited exposure.

The operative lesson analysts drew: discovery is now the easy part; verification and remediation throughput are the bottleneck, which is exactly the demand ShipItClean monetises.

The accessibility contrast. Mythos is widely regarded as the most powerful AI model on the planet and is deliberately access-restricted to a limited group of critical industry partners; Mozilla's result required a custom pipeline run over roughly a month. ShipItClean scanned the same Firefox repository with a commodity 14-billion-parameter model on Atlas for about $15 of electricity. The frontier tool is locked to a handful of partners. The ShipItClean scan is available to anyone, at a price two orders of magnitude smaller. Internal

Competitive structure

ShipItClean's published comparison pages position it as complementary to every named competitor and recommend running both. That reads at first like conflict avoidance. It is better understood as the correct reading of the structure: the tools genuinely do different units of work, and the honest positioning is also the one that survives platform consolidation.

There is a structural point beneath the positioning one. The AI-native reviewers, CodeRabbit, Greptile, Qodo, Copilot review, are essentially API wrappers: they orchestrate third-party frontier models and therefore carry per-token inference costs they cannot control, marked up inside flat seat pricing. Their gross margin is set by their model vendor's price list, and every price increase from that vendor compresses it. ShipItClean, moving onto an owned model, can price its scans below what wrapper competitors pay their vendor per token and still hold a high margin. That is not a feature advantage; it is a cost-structure moat that no wrapper can match without vertically integrating a model of its own. Internal

Positioning by scope of analysis and cost-base control A two-axis map. The horizontal axis runs from vendor-controlled inference cost on the left to owned inference on the right. The vertical axis runs from pull-request diff scope at the bottom to whole-codebase scope at the top. AI reviewers cluster lower left, rule engines sit mid right, and ShipItClean sits alone in the upper right corner where whole-codebase scope meets owned inference. COST BASE CONTROL --> vendor API owned inference SCOPE --> whole codebase PR diff UNCONTESTED CodeRabbit Greptile Qodo Copilot review Semgrep Sonar Snyk ShipItClean ShipIt
The vertical axis is what gets analysed: the diff, or everything. The horizontal axis is who controls the inference bill. The reviewers cluster where scope is bounded by a context window and margin is set by a vendor price list. The rule engines own their economics but are bounded by what a pattern can express. The upper-right quadrant is empty, and the Firefox audit is the evidence that it is reachable.

Why the upper-right quadrant is empty

Not because nobody wants it, but because a context window makes it structurally impossible. No model, frontier-scale included, can natively reason across 1.5 billion tokens of interrelated code. Greptile's codebase graph and Copilot's March 2026 agentic context-gathering are both real improvements, and both still retrieve into a bounded window.

Atlas has no fixed context ceiling; its capacity is bounded by memory and engineering rather than a token limit. The Firefox audit is the proof, and its design is deliberately unflattering to the model: a 14-billion-parameter model, chosen specifically because it has no frontier reasoning ability, held 44M tokens of source and ran 16,381 consensus passes. Internal

To make the scale concrete: Microsoft Windows is plausibly the largest single codebase on the planet, roughly 3.5 million files and 270-300 GB, requiring Microsoft to build a virtual file system just to let 4,000 engineers work in it. Chunk-bound reviewers cannot reason about a system like that at all; they see keyhole after keyhole. Atlas could hold it in memory whole, letting a model trace cross-module data flows, trust boundaries and chained defects that only appear when the entire system is in view. Internal

The argument this makes

Beating a frontier model was never the point of the Firefox run. Isolating the variable was. A small model backed by Atlas did what no model can do unaided, which establishes that the barrier to whole-system code reasoning is architectural, not a matter of model size. Atlas is the multiplier; the model is a choice. Put a frontier model behind it and the same result scales.

The bundling question, answered by MCP

Platform consolidation is real. Graphite, last valued at $290M, was acquired by Cursor in December 2025, and GitHub shipped agentic Copilot review in March 2026. Cited For a per-seat PR reviewer this is close to an extinction event.

ShipItClean exposes its scan over MCP, so any AI coding agent or CI pipeline can call it natively with no bespoke integration. Internal That inverts the threat: when Copilot's or Cursor's agent calls the scan, consolidation of the coding surface becomes a distribution channel rather than a competitor. The bundlers are annexing the review layer. They are not building whole-codebase adversarial audit, and the context ceiling is why they cannot.

The bundle is not free either. Copilot review stopped being included in the base subscription from 1 June, billing token use in AI Credits, with agentic infrastructure consuming Actions minutes on private repositories. Cited Its advantage is default placement, not price.

Four differentiators that are already shipped

These are not roadmap items. Each is live, and each is independently checkable by a prospect.

Cost base

The five-tier routing ladder runs from Claude-class inference at $0.0140 per thousand tokens to the owned 14B model at $0.0002, which is electricity only. That places marginal cost roughly 70x below a Claude-class competitor and 35x below a DeepSeek-class one. Internal

Current boundary. The Firefox audit demonstrates these economics at production scale, so the cost advantage is demonstrated rather than theoretical. The public product currently routes to third-party APIs; only internal accounts run on the local model. Migrating public traffic onto the owned model is simultaneously the largest value-creation event available and the principal execution risk.

One current-week data point sharpens this. On 6 August 2026, DeepSeek, the low-cost anchor whose loss-leader pricing had held the entire API market down and forced ByteDance, Tencent and Western vendors to cut rates, warned developers of a significant, imminent price increase (founder Jun Song indicated 2x to 10x), explicitly ending its price-war posture. Cited DeepSeek's below-cost pricing had been the anchor that kept the whole market's token prices low. With that anchor lifting, every provider above it gains room to raise prices as competitive pressure eases; providers who held rates down to match DeepSeek will most likely follow. That increase flows straight into the COGS of every API-wrapper competitor, compressing margins or forcing seat-price increases at exactly the moment ShipItClean's owned-model cost base stays flat. The relative advantage widens rather than erodes.

The free tier as a wedge

The free deterministic scan already delivers the substance of what an entire class of competitors sells as a paid product: the proprietary Hyrex pattern engine alongside Gitleaks for secrets, Trivy for dependency CVEs against NVD and GitHub Advisory, IaC misconfiguration checks, Semgrep with 2,000+ rules across 30+ languages, Bandit, ESLint and flake8, with AI synthesis grouping the findings. That covers for free what GitGuardian, Snyk Code, SonarQube and Mend charge for. Internal

The paid layer is where the moat sits: 108 adversarial personas run over the whole codebase in Atlas, finding logic errors, design flaws and multi-step attack chains that have no rule yet.

The flywheel

Hyrex is built from the paid AI scans. Every paid engagement sharpens the free tier's rule base. That is a compounding data asset that static rule vendors structurally cannot build, and it converts the free tier from a customer-acquisition cost into a product that improves as revenue grows.

Uncorrelated reviewers

Agents are routed across different model families, including Claude, DeepSeek and GPT-4o, specifically to avoid correlated blind spots, rather than running many prompts against a single model. Internal On the Firefox audit the consensus layer distilled 36,317 raw findings to 72 confirmed threats, filtering out the 90%-plus that were test code, detection code, documentation and architecture opinions.

That filtering ratio is the commercially important number, not the raw finding count. Every AppSec buyer's actual complaint is false-positive volume.

Remediation, not just detection

Nearly every competitor stops at the finding. Every completed scan here emits two artifacts: a Download Fix Workflow containing the actual patched source files with a change log and a verify-then-rescan checklist, and a Copy Fix Workflow structured for pasting into a coding agent. Internal

Two design choices make the second one work where competitors' equivalents do not. It is pre-filtered through the paid consensus layer, so the agent is told to fix only genuine threats rather than chase noise into regressions. And it instructs the agent to work one issue at a time in order rather than batching, because an agent asked to fix everything at once exhausts its context across files and produces sloppy edits.

The next model: purpose-built for adversarial audit

The current product routes scans through third-party frontier models. The owned model demonstrated on Firefox proved the cost structure. The next step is not to run a general-purpose model more cheaply. It is to build a model whose architecture is designed around the audit task itself.

B²: triple-context architecture

B² splits the decoder into two halves: a frozen self-decoder that retains general language and code comprehension, and a trained cross-decoder that grounds every generation in three independent context streams. One stream holds scanning rules, vulnerability taxonomies and exploit patterns. Another holds the full codebase via Atlas. The third holds the active audit reasoning. The architecture means the model does not choose between knowing code and knowing security -- it holds both simultaneously in separate, non-competing contexts.

A general-purpose model repurposed for security review is a compromise: it allocates the same attention budget to audit reasoning that it allocates to writing poetry. A model whose cross-decoder was trained exclusively on adversarial code analysis is not. Every parameter in the trained half exists to find defects.

What this changes

The owned-model migration is not a cost-reduction exercise that risks quality. It is an architecture upgrade: a model purpose-trained for the exact task, running on owned infrastructure, at a cost base no API wrapper can reach. The moat is not just cheaper inference. It is a model that gets better at security audit specifically, compounding on every scan it runs.

The architecture is not theoretical. A 7B² demo has been implemented, trained and released as an open-weight model that anyone can download, run and verify: Sharona-B2-7B-Jbliterated on Hugging Face. This is not a roadmap slide. The triple-context architecture works, the training pipeline works, and the next step is a security-specialized variant trained on the corpus of every scan ShipItClean has run.

Proof points

Independently checkable results
ResultWhat it demonstrates
Firefox: 44M tokens of source, 36 of a 108-agent roster, ~1.6B effective tokens, 16,381 consensus passes, ~$15 electricity against ~$50,000 quoted; 36,317 findings distilled to 72 confirmed. Full report is public. InternalThe context ceiling is architectural, and the cost structure holds at scale
Enterprise Linux installer: 54 vulnerabilities found, including live plaintext credentials for private repositories CitedFinds real, exploitable, high-consequence defects in shipped third-party software
Self-scan: 158 findings, 3 critical, all fixed CitedThe tool is run against its own codebase and the results are published
Mozilla / Mythos: Anthropic's frontier model found 271 vulnerabilities in Firefox 150 (180 sec-high, 80 sec-moderate, 11 sec-low) via a bespoke month-long pipeline CitedIndependent confirmation that whole-codebase audit finds real defects in hardened code; ShipItClean scanned the same repo at commodity scale and cost

The Firefox result is the one to lead with, because it is the only one that simultaneously proves the technical claim, the cost claim and the false-positive claim in a single artifact a prospect can verify without trusting anyone.

Risks

Owned-model migration. The cost moat is demonstrated but not yet carrying public traffic. This is the central technical bet and the largest single risk. Internal

Category legibility. Selling a unit of work the market does not yet have a budget line for is harder than selling a cheaper version of something familiar. Buyers have a line item for per-seat review and a line item for pentest engagements; whole-codebase continuous adversarial audit sits between them. The complementary positioning mitigates this by attaching to budgets that already exist, but it does not eliminate it.

Benchmark unverifiability. Competitor quality claims are vendor-run and mutually inconsistent. Greptile publishes 82% bug-catch, Qodo 60.1% F1 on its own test. Cited Buyers cannot verify any of them, including ShipItClean's. ShipItClean's published full reports are the strongest response available, because they are the only claim in the category that survives independent scrutiny.

Incumbent response. Snyk, Sonar and GitHub have distribution and could fund an Atlas-equivalent. Defensibility rests on the architecture and cost structure rather than on features. Internal

Inference price direction. A renewed race toward near-free third-party inference compresses the cost advantage, though the context ceiling would persist regardless. Internal

What is missing

Four gaps limit how far this document can be sent. The first three are the inputs to an obtainable-market figure. The fourth is the one that gates the entire technical argument.

  • Enterprise price per scan and audit cadence. Without these there is no revenue model to underwrite.
  • Count of organisations in the stated SAM. Large, sensitive or regulated codebases needs a number attached.
  • Operating metrics. Revenue, customer count, retention, and free-scan-to-paid conversion. The entire product-led motion rests on that last one.
  • Owned-model audit quality benchmarked against the routed tiers. The most important missing number in the document. The Firefox run proves Atlas removes the context ceiling. It does not establish that the 14B model's findings match what a frontier model behind Atlas would produce, and until that comparison exists the migration risk cannot be fully quantified.

Sources

  1. ideaplan: AI Code Review Tools Market Share 2026
  2. Grand View Research: AI Code Tools Market Size and Share Report
  3. Mordor Intelligence: AI Code Tools Market Size, Share and Growth Trends
  4. MarketsandMarkets: Application Security Testing Market Report 2025-2030
  5. Business Research Insights: SAST Tool Market Size and Trends
  6. Grand View Research: SAST Application Security Market Statistics
  7. Keyhole Software: Software development statistics 2026
  8. CloudBees: 2026 State of Code Abundance Report
  9. Larridin: AI code share, empirical study of 4.2M developers
  10. Veracode: Spring 2026 GenAI Code Security Update
  11. AppSec Santa: DevSecOps statistics 2026
  12. Cloud Security Alliance: AI-generated code vulnerability surge, 2026
  13. DEVOPSdigest: IDC InfoBrief, security cost per developer
  14. Elisity: 2026 cybersecurity budget benchmarks
  15. Particula Tech: Greptile vs CodeRabbit vs Qodo, 2026 (Graphite acquisition, Qodo benchmark)
  16. Tech Funding News: Greptile valuation and bug-catch claim
  17. Codacy: GitHub Copilot code review billing change
  18. Developers Digest: Agentic Copilot review, March 2026
  19. CodeRabbit: CodeRabbit pricing 2026
  20. ShipItClean: Published scan reports and investor summary
  21. Microsoft: Copilot 20M+ users, 90%+ of Fortune 100, ~400% YoY growth (Jul 2025 earnings)
  22. NYU: "Asleep at the Keyboard? Assessing the Security of Code Contributions from Large Language Models" (IEEE S&P 2022, arXiv 2108.09293)
  23. IBM: Cost of a Data Breach 2025, shadow AI supplement (~$670K additional breach cost)
  24. Mozilla / Anthropic: 271 AI-found Firefox vulnerabilities fixed in Firefox 150 (Apr-May 2026); SecurityWeek, Help Net Security, The Decoder, Schneier on Security; Mozilla technical writeup (May 7, 2026)
  25. DeepSeek: API price-hike notice, 6 August 2026 (SCMP, TechNode, Dataconomy)
ShipItClean is powered by our CodeForge Engine Ask AI About Us
Privacy Policy  ·  Terms of Service  ·  AI Overview
S
Sharona-AI
Online