All news
aiproduct

GPT-5.6 Vision Lineup: Sol, Terra, Luna Benchmarked

31 Aug 2026

OpenAI has released its GPT-5.6 lineup, introducing three new vision-capable models — Sol, Terra, and Luna — via a release stream that showcased UI agents and detailed 3D visualizations. Early benchmark comparisons against GPT-5.5 and third-party models paint a mixed picture: strong gains in object detection and counting, but notable weaknesses in OCR, cost, and stability at high resolutions.

What the benchmarks show

Across object detection (mAP@50), Sol posted a sizable jump over its predecessor:

  • Object detection (mAP@50): GPT-5.5 13.8 | Sol 46.2 | Terra 44.7 | Luna 43.3
  • Object counting: GPT-5.5 64.9% | Sol 73.0% | Terra 67.6% | Luna 66.2%
  • OCR mean similarity: GPT-5.5 91.2% | Sol 90.7% | Terra 88.8% | Luna 88.4%
  • Text extraction: GPT-5.5 87.6% | Sol 82.5% | Luna 81.4% | Terra 79.4%
  • Average speed per image: Sol ~10s | Terra ~6s | Luna ~5s
  • Average cost per image: Sol ~2.5¢ | Terra ~1¢ | Luna <0.5¢ | Gemini 3.5 Flash 0.8¢

Notably, Claude Fable 5 was flagged as the most expensive model in the comparison, edging out Sol on a per-image cost basis, while Gemini 3.5 Flash undercuts Sol on price while leading in both detection and counting metrics.

A broader public vision benchmark comparing these models is expected within the next few weeks, which should give founders more independently verifiable data.

Risks and caveats

The report flags several implementation-sensitive issues:

  • Coordinate format sensitivity: Using the wrong coordinate format can drop Sol's detection performance by roughly 15 mAP points — a significant swing tied purely to implementation details rather than model capability.
  • Large-image instability: Sol becomes less stable on images around 2,000x2,000 pixels or larger, especially at lower reasoning-effort settings.
  • OCR/text extraction lag: Sol underperforms GPT-5.5 in both OCR mean similarity and text extraction, raising questions about whether the "best vision model" framing holds for document-heavy workloads.
  • Cost and speed tradeoffs: Sol is more expensive and slower per image than Terra and Luna, while Gemini 3.5 Flash — a competing model — offers lower cost and leading detection/counting scores.

Sources do not specify the exact datasets or methodology behind these benchmarks, nor do they explain why Sol underperforms GPT-5.5 on OCR despite leading elsewhere. Model architecture and pricing-tier details for Sol, Terra, and Luna were also not disclosed.

Why founders should care

If your product relies heavily on object detection or counting, Sol's benchmark gains suggest it's likely worth testing — the jump from GPT-5.5's 13.8 to Sol's 46.2 mAP@50 is substantial, though real-world performance will probably depend on correctly implementing coordinate formats.

For budget- or latency-sensitive applications, Terra and Luna are plausibly better starting points given their lower cost and faster per-image speed, especially if top-tier detection accuracy isn't mission-critical.

For document processing or OCR-centric products, founders should be cautious: switching to Sol may not deliver improvements and could even underperform GPT-5.5 on these specific tasks.

Given that Gemini 3.5 Flash reportedly beats Sol on both cost and detection/counting metrics, founders evaluating vision models should probably benchmark it directly rather than assuming OpenAI's newest release is the default best option.

Finally, teams working with high-resolution imagery (2,000x2,000 pixels or larger) should test Sol at their target resolution and reasoning-effort setting before deployment, given the reported instability at scale.

Bottom line

GPT-5.6's Sol model shows real improvement in detection and counting tasks, but it's not a universal upgrade — cost, speed, OCR accuracy, and large-image stability all cut against a blanket "best vision model" claim. Founders should treat these benchmarks as a starting point for their own testing, particularly once the promised public vision benchmark becomes available in the coming weeks.

Sources