Kimi K3 Tops New Benchmark of AI Models for Geological Reasoning

Model scores across three datasets: Kimi K3 (88.1%), Sonnet 5 (86.2%), GPT 5.6 Sol (85%), Ox Alpha/GLM 5.3 Flash (84.4%), Deepseek V4 Pro (77.9%), Grok 4.6 (77.3%), Gemini 3.1 Pro Preview (75.5%), and GLM 4.7 (6.81).

Relative rankings of AI models on the Groundtruth benchmark

Breakdown of Groundtruth benchmark leaderboard scores across USGS, Wamex and Technical datasets. Adjacent gaps, paired permutation test on all 150 questions: Kimi → Claude +0.19 (p=0.1111, tied) · Claude → GPT +0.12 (p=0.3866, tied) · GPT → OxAlpha +0.06

Breakdown of Groundtruth benchmark leaderboard scores

Spread of per-question token spend. Highest cost: Gemini 3.1 Pro, lowest cost: Ox Alpha (free), followed by Deepseek V4 Pro.

Cost per question answered by each leading model

Groundtruth benchmark tests leading AI models on questions generated from real-world geological datasets

Rather than asking a standard set of questions, Groundtruth generates a fresh set from your own dataset, letting you compare models and agent harnesses on the data that matters to you.”
— Jen Dodgson, Eigenform CEO
SINGAPORE, SINGAPORE, August 28, 2026 /EINPresswire.com/ -- Kimi K3 recorded the highest overall score on the new Groundtruth benchmark, designed to test how well leading AI models can reason over novel real-world earth science information.

The Groundtruth benchmark tested models from OpenAI, Anthropic, Deepseek, Google, XAI and Moonshot, as well as the new stealth release OxAlpha (GLM 5.3 Flash), across three geological datasets spanning historical exploration records, technical commodity reports and academic geology. Rather than relying on a fixed bank of general-knowledge questions, Groundtruth generates new evaluations from geoscience source material, allowing models to be tested against unfamiliar datasets and region-specific data.

Kimi K3 achieved the highest aggregate score of 91%, closely followed by Claude and GPT instances. However, confidence intervals for the leading models overlapped, placing several systems in the same performance tier. The evaluation also found substantial differences in inference cost between models achieving comparable scores. Kimi, Claude, GPT and Deepseek offered broadly similar price-per-correct-answer rates, but were outcompeted on price by Ox Alpha, which is currently available free.

The results highlight a growing problem in evaluating AI for specialist scientific work. General-purpose benchmarks can measure broad reasoning ability, but provide limited evidence of how a model will perform when confronted with the reports, terminology, historical interpretations and incomplete evidence characteristic of a particular technical domain.

Groundtruth was developed by Singapore-based AI research company Eigenform AI, using their background in recursive self-improvement within complex dynamic environments to build a responsive alternative to existing benchmarks. The goal was to build a framework that would allow industry and academic users to create custom evaluations based on their own data to test alternative models, retrieval systems and harnesses.

Eigenform CEO Jennifer Dodgson said: “If you’re an exploration manager with 15 years of reports, logs, maps, assays and consultant studies sitting in SharePoint and a vendor wants you to buy their AI system they will tell you it scored 83% on some general academic benchmark, but that’s not important to you. You don’t want a model that knows bismuth can be a pathfinder for gold, you want a model that can work out whether it serves as a pathfinder in your specific terrain. That’s why we made the Groundtruth Dynamic Benchmark. Rather than asking a standard set of questions to test the sector knowledge of a model, Groundtruth generates a fresh set of questions and scoring rubrics from your own dataset, letting you compare models and agent harnesses on the data that matters to you.”

To create a representative leaderboard Eigenform worked with tenement mapping and marketplace service, NextMaps, and commodities analysis platform MatchPoint to create test datasets covering a variety of regions and stages in the exploration process.

MatchPoint CEO, Manessa Mungroo said: “AI is becoming an essential part of the resources value chain, and transparency is vital to its effective adoption. At MatchPoint, we have fully incorporated LLMs into our data extraction algorithms. Large volumes of unstructured data are now being processed into structured formats at scale, enabling discoveries and insights that would previously have taken days for an analyst to generate.”

Akira Rafhael
Eigenform Pte Ltd
+65 9133 0611
alex@send.eigenform.ai
Visit us on social media:
LinkedIn
X

Legal Disclaimer:

EIN Presswire provides this news content "as is" without warranty of any kind. We do not accept any responsibility or liability for the accuracy, content, images, videos, licenses, completeness, legality, or reliability of the information contained in this article. If you have any complaints or copyright issues related to this article, kindly contact the author above.