Submersion AI's Basin enters CyberGym top 10 at launch
Submersion AI has launched Basin, a specialist AI model for cybersecurity reasoning that debuted at 8th place on the CyberGym benchmark, the widely cited leaderboard for AI-assisted vulnerability discovery. The New York-based company says Basin scored 80.8% on the benchmark at launch, surpassing Grok 4.7 (80.3%) and Anthropic's Opus 4.8 (78.8%), and sitting close behind Mythos at 83.2%, which currently holds a higher position on the table.
The company positions Basin as a smaller, lower-cost alternative to frontier general-purpose models adapted for security use, claiming it is up to 90% more cost-effective than comparably performing systems. Submersion AI did not publish full parameter counts or hardware requirements in its release.
Real-world vulnerability discovery
Basin has already been used to uncover two publicly disclosed, high-severity vulnerabilities. The first, CVE-2026-86733, affects Snipe-IT before version 8.7.0 and allows an authenticated superadministrator to trigger remote code execution by uploading a malicious ZIP backup through the restore function. The second, CVE-2026-82450, affects BookStack before version 26.05.4, where an authorised user can achieve unauthenticated remote code execution by uploading a portable ZIP import containing a PHP polyglot file as a book cover, bypassing extension validation and executing in the public web root. Both CVEs are credited to the Submersion AI Security Research Team.
The model's deployment architecture is designed with data-residency constraints in mind. Basin can run fully within a customer-controlled environment, whether on-premises, in a private cloud, or air-gapped, addressing the usage-policy restrictions that prevent mainstream hosted AI platforms from being used in sensitive security-testing workflows. A hosted API is also available for teams without those constraints. The company says Basin performs autonomous reconnaissance, probes applications, chains findings into validated exploit chains, and produces analyst-ready reports.
Market context
The market for AI-assisted penetration testing and vulnerability research is expanding rapidly, driven both by an acute shortage of skilled offensive-security professionals and by the growing complexity of enterprise attack surfaces. Established players in the automated security-testing space include Synack, HackerOne, and a range of vulnerability-scanning vendors, while larger AI labs have begun offering security-specific model variants, often with usage restrictions that limit their utility in authorised red-team or research contexts.
Submersion AI's emphasis on air-gapped and on-premises deployment addresses a genuine gap: government agencies, defence contractors, and critical-infrastructure operators routinely handle systems that cannot connect to external inference APIs, and they face regulatory obligations under frameworks such as FedRAMP, NIST SP 800-53, and, in Europe, NIS2 and DORA. The ability to deploy a frontier-grade reasoning model entirely within a controlled perimeter removes a structural barrier that has kept cloud-hosted AI tools off the most sensitive networks.
Standards and regulatory read-across
The emergence of capable, deployable offensive-security AI models will inevitably attract policy scrutiny. The EU AI Act classifies systems used in critical infrastructure and law enforcement as high-risk, and AI tools capable of autonomous exploitation could fall under additional obligations depending on how regulators interpret the "safety component" provisions. In the US, the administration's AI executive orders and CISA's evolving guidance on AI in cybersecurity are likely to shape procurement decisions for any federal or defence-adjacent customer that Submersion AI targets.
The company was founded by practitioners from offensive security, national security, and vulnerability research backgrounds, and has not disclosed funding, headcount, or named enterprise customers. Near-term milestones to watch include independent replication of the CyberGym benchmark results, additional CVE credits, and any named commercial or government deployments.