All news
aiproductsaas

Bullshit Detector & hwatu: New Open-Source Agent Tools

30 Jul 2026

Two new agent tools, two different problems

This week's Show HN queue produced two open-source projects aimed squarely at builders working with AI agents: Bullshit Detector, an agent skill that fact-checks videos and articles, and hwatu, a lightweight verification browser designed to let AI agents check web pages far faster than traditional browser automation.

Neither tool has adoption numbers, funding, or a long track record yet — both are early-stage open-source releases. But the mechanics and benchmarks reported are specific enough to be worth a look for founders building in the AI-agent tooling space.

Bullshit Detector: claim-by-claim verification for content

Bullshit Detector is a "Portable Agent Skill" — built in plain markdown and self-contained Python — that fact-checks YouTube videos, articles, tweets, and PDFs. It extracts claims, verifies them against independent sources, and outputs a claim-by-claim report with a BS score from 0–10.

The workflow includes:

  • A hype-signal scan
  • Incentive analysis (who benefits if a claim is believed)
  • Per-claim verdicts that require cited sources
  • A fetch-content script pulling YouTube transcripts and TikTok captions via yt-dlp, plus article and PDF extraction and tweet retrieval — all without API keys
  • A Whisper-based fallback (mlx-whisper on Apple Silicon) for transcribing caption-less TikToks and Reels, with no system ffmpeg required

In testing cited in the release, a "make money with AI" video (1.16M views) scored 5/10, with 12 claims broken down into 4 confirmed, 2 plausible, 3 misleading, 0 false, and 3 unverifiable. A TikTok claiming "our Sun has a hidden twin" (552K views) scored 9/10. Notably, the tool's own README self-scored 3/10 when run through the detector itself — a useful (if humbling) sanity check on how the scoring behaves.

The project is open source under the MIT license and is compatible with Claude Code, Codex, OpenCode, and other harnesses supporting the skills format. It was reportedly built using Claude Code itself.

hwatu: a faster verification browser for agents

hwatu targets a different problem: giving AI agents a fast way to check whether a rendered web page looks and behaves as expected, without the overhead of spinning up a full browser automation stack.

Reported performance figures:

  • Window creation: 13 ms
  • Check command: ~35 ms median, reportedly 9x faster than warm-server Playwright
  • Cold start: 435 ms (versus Playwright's 190 ms — the report doesn't explain this gap despite hwatu's overall speed claims)
  • A full Playwright verification pass: 82 ms best case with five API calls; a pass shaped like hwatu's runs takes 341 ms versus hwatu's 39 ms
  • An agent using hwatu improved pixel match from 85.1% to 98.8% on a cloned Stripe landing page

hwatu ships as a single static binary plus the distro's WebKitGTK library, and can be used via an MCP server, CLI, or a 1-line JSON socket protocol. A "focus" command can materialize a live session in a tiling window manager for human intervention. It supports Claude Code, Cursor, Jcode, and generic MCP workflows, and can be installed via curl script, Arch's package manager, or built from source with cargo.

The catch: hwatu is Linux-only today, renders with WebKit rather than Chromium (meaning teams may still need a separate Playwright matrix in CI to catch Chromium-specific bugs), and is licensed under AGPL-3.0, which carries copyleft obligations that could complicate use in proprietary products — a contrast to Bullshit Detector's more permissive MIT license.

What's missing

The report is upfront about several gaps: there's no independent validation of Bullshit Detector's BS scores against human fact-checkers, no adoption or download data for either tool, no pricing or monetization details, and no explanation for hwatu's slower cold-start time despite its faster steady-state performance. There's also no stated connection between the two projects beyond both surfacing as agent-tooling launches around the same time.

Why founders should care

  • Founders building AI-agent products may increasingly need built-in verification layers — for content accuracy or for UI/QA — and these two projects suggest that demand is starting to attract dedicated tooling, though it's too early to say either will become a standard.
  • hwatu's benchmarks, if they hold up outside the reported test cases, suggest agent-based QA pipelines could plausibly run faster and cheaper than Playwright-based setups — a potential cost lever for teams running frequent automated checks, though this remains unverified at scale.
  • The licensing split (MIT vs. AGPL-3.0) is a practical consideration: teams evaluating hwatu for commercial products should weigh the copyleft implications before integrating it into proprietary codebases.
  • Both tools' compatibility with Claude Code, Cursor, and MCP-based workflows indicates they could likely be trialed with minimal integration effort inside existing AI-agent stacks — a low-cost way to test whether either delivers real value before committing.
  • The use of Claude Code to help build Bullshit Detector hints at a broader pattern: solo builders and small teams are likely shipping increasingly specific developer tools faster with AI coding assistants, which may lower the bar for founders considering similar niche tooling plays.

As with any early-stage open-source release, reliability at scale, security posture, and long-term maintenance remain open questions. Founders evaluating either tool should treat the current benchmarks as promising signals rather than proven guarantees.

Sources