AIThis post was created with the assistance of artificial intelligence (AI).
VigilSAR Defense LLM Benchmark
The public benchmark page — aggregate results public, task set private. Source: vigilsar.com

VigilSAR, a defense-ISR software product, has published an LLM benchmark focused on whether language models can be trusted with intelligence-surveillance-reconnaissance work. It measures the reasoning, reporting, and restraint an analyst actually needs—not performance on general trivia.

The evaluation covers 14 models across 300 tasks, scored on 2026-07-17. Aggregate results appear on the public leaderboard, giving tech readers a comparative view without exposing the underlying evaluation material.

The task set is deliberately private so models cannot train on it. A separate private held-out set adds another check, while the published gap between public and held-out scores for each model helps flag memorization.

In the current standings, claude-fable-5 leads with 67.77 in Band A and serves as the pinned reference row. The benchmark emphasizes bands rather than rank because confidence intervals within a band overlap.

VigilSAR public LLM leaderboard
The leaderboard — compare bands, not rank numbers. Source: vigilsar.com/benchmark

The notable new entry is Moonshot’s Kimi K3, debuting at #3 with 64.65 in Band B. That places it ahead of every GPT and Gemini row on the board.

The GPT-5.x family fills Bands C-D, while Gemini rows sit in Bands E-F. One locally runnable open model is scored as “sovereign-deployable”, reflecting that deployment reality is part of the score.

The benchmark starts from a blunt premise: “Vendor claims are not evidence.” Its operators built the evaluation to decide which models get anywhere near their own product and to rank the models they use themselves. They state that they are not paid by any vendor and that “we would rather be measured than believed.”

Its honesty features reinforce that position: bands instead of pseudo-precise ranks, published confidence intervals, published held-out gaps, and a pinned reference row. The board also reports per-model cost-per-correct-answer economics, pairing capability results with practical model economics.

Powered by Thorsten Meyer AI


Engineering with Small Language Models: Efficient AI Design, Training, and Deployment for Developers

Engineering with Small Language Models: Efficient AI Design, Training, and Deployment for Developers

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Cyber Reconnaissance, Surveillance and Defense

Cyber Reconnaissance, Surveillance and Defense

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

AI Engineering: Building Applications with Foundation Models

AI Engineering: Building Applications with Foundation Models

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The Emerging Science of Machine Learning Benchmarks

The Emerging Science of Machine Learning Benchmarks

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

You May Also Like

Continuous Upskilling: Using AI to Identify and Fill Skill Gaps in Your Team

Keeping your team’s skills sharp requires AI insights that uncover hidden gaps—discover how to unlock your team’s full potential today.

Is Artificial Intelligence the Next Step in Animal Communication?

Just when we thought we understood animals, AI may unlock a new realm of communication—find out how it’s changing everything.

AI Augmentation Success: Stories of AI Making Humans More Effective

AIThis post was created with the assistance of artificial intelligence (AI).AI augmentation…

AI Co-Workers: How Teams Are Collaborating With AI Tools Daily

Gaining insights into how teams are seamlessly integrating AI co-workers reveals transformative collaboration methods that could reshape your workplace.