Sqlsure: Deterministic Checks for AI-Generated SQL
12 Jul 2026
A new gatekeeper for AI-generated SQL
As more startups ship products that let AI agents write and run SQL queries, a persistent problem has emerged: LLMs can generate syntactically valid SQL that is semantically wrong — the kind of error that silently double-counts revenue or joins the wrong tables. A newly released open-source tool, sqlsure, aims to catch these errors before they run.
Announced via Show HN and installable via pip, sqlsure performs deterministic semantic checks on AI-generated SQL queries. Unlike approaches that rely on another model to review output, sqlsure uses rule-based checks that execute in roughly 0.1 milliseconds before a query is allowed to run.
What the testing showed
The tool's creators tested sqlsure against 2,568 expert-written queries drawn from two (unnamed) text-to-SQL benchmarks. Results reported include:
- 45 flags identified across the benchmark queries, with zero false alarms
- 10 out of 10 fixes applied verbatim to flagged queries produced a passing result
- 2 foreign keys recovered that were missing from the BIRD benchmark's published schema
- 16 out of 16 rule tests achieved 100% recall and 0% false positives when paired with a benchmark
The report does not specify which two benchmarks were used beyond referencing BIRD by name, nor does it detail exactly which classes of semantic errors sqlsure is designed to catch.
How it's meant to be used
Sqlsure is positioned as flexible infrastructure rather than a single-purpose app. According to the report, it can be deployed in three ways:
- As a CI gate — blocking merges when generated SQL would, for example, double-count revenue
- As an MCP server — requiring AI agents to pass inspection before executing any query
- As an embedded library — integrated directly into text-to-SQL products and agent frameworks
This multi-mode deployment suggests the tool is aimed squarely at teams building AI data agents rather than end users querying databases directly.
Why founders should care
For founders building text-to-SQL products, AI data agents, or internal analytics tools powered by LLMs, sqlsure's benchmark results are a signal worth watching — though the sample size and lack of production-scale validation mean the numbers should be treated as promising rather than conclusive.
- The sub-millisecond execution time suggests it's plausible sqlsure could be added to real-time pipelines without noticeable latency, though this is likely to vary with schema complexity in production.
- Zero false alarms across 45 flags in testing could indicate a low-friction path to CI/CD adoption, but broader validation across more diverse, production-scale workloads would likely be needed before founders can be confident the false-alarm rate holds at scale.
- The tool's ability to surface missing foreign keys hints that it may double as a schema-documentation check — potentially useful for teams that already suspect their schema metadata is incomplete.
- Because it's available via pip with three integration modes, founders may find it relatively easy to trial without committing to a specific architecture upfront.
- More broadly, the release signals that demand for verification layers around LLM-generated SQL is likely growing as more startups put AI agents in charge of production data queries — a trend founders shipping AI data tools should probably track closely.
What's still unclear
The report leaves several open questions: there's no information on who built sqlsure or what organization stands behind it, no licensing or pricing details, no comparison to existing SQL validation tools, and no data yet on community adoption or supported SQL dialects. Founders evaluating the tool for production use will likely want to wait for more of these details — or test it directly against their own schemas — before relying on it as a sole safeguard.
Risks to weigh: deterministic, rule-based checks may not catch every class of semantic error, especially novel patterns not represented in the benchmark queries tested. The benchmark itself, while sizable at 2,568 queries, may not fully represent the schema complexity of real production workloads.