Kimi K3 leads new geology AI benchmark
Kimi K3 posted the top score on Groundtruth, a new benchmark built by Singapore-based Eigenform AI to test how AI models reason over real geological data. The results also showed that several leading models performed in the same tier, while cost differences could matter as much as raw accuracy.
Why it matters: - Groundtruth tests AI models on unfamiliar, domain-specific earth science material instead of generic trivia. - The benchmark is meant to show how models perform on the kinds of reports, terminology, historical interpretations and incomplete evidence used in exploration work. - The results suggest raw benchmark scores can miss major differences in cost and usefulness for specialist scientific tasks.
What happened: - Kimi K3 recorded the highest overall score on Groundtruth with 91%. - The benchmark compared models from OpenAI, Anthropic, Deepseek, Google, xAI and Moonshot. - Groundtruth also tested the stealth release OxAlpha, listed as GLM 5.3 Flash. - The evaluation used three geological datasets covering historical exploration records, technical commodity reports and academic geology. - Groundtruth generated new questions from geoscience source material rather than using a fixed question bank. - The benchmark was developed by Singapore-based Eigenform AI.
The details: - Eigenform said the framework is designed so industry and academic users can build custom evaluations from their own data. - The benchmark is intended to test models, retrieval systems and agent harnesses on region-specific and proprietary material. - Confidence intervals for the leading models overlapped, placing several systems in the same performance tier. - Kimi, Claude, GPT and Deepseek showed broadly similar price-per-correct-answer rates. - Ox Alpha was cheaper on price and is currently available free. - Eigenform worked with tenement mapping and marketplace service NextMaps and commodities analysis platform MatchPoint to build representative test datasets. - The test sets covered a range of regions and stages in the exploration process. - MatchPoint has incorporated large language models into its data extraction algorithms. - MatchPoint said its systems process large volumes of unstructured data into structured formats at scale.
Between the lines: - General-purpose AI benchmarks can show broad reasoning ability, but they may not predict performance on specialized technical work. - Groundtruth points to a shift toward dynamic, data-specific evaluation tools for scientific and industrial users. - The overlap among top scores suggests the practical gap between leading models may be smaller than headline rankings imply.
What's next: - Eigenform is positioning Groundtruth as a customizable benchmark for organizations that want to test AI on their own datasets. - More users in exploration, geology and adjacent technical fields may adopt private evaluations before buying or deploying AI systems. - Cost-per-answer pressure is likely to become a bigger factor in model selection as more systems reach similar accuracy levels.
The bottom line: - Kimi K3 topped a benchmark built to measure real-world geological reasoning, but the broader takeaway is that domain-specific evaluation and pricing may matter more than a single score.
Disclaimer: This article was produced by AGP Wire with the assistance of artificial intelligence based on original source content and has been refined to improve clarity, structure, and readability. This content is provided on an “as is” basis. While care has been taken in its preparation, it may contain inaccuracies or omissions, and readers should consult the original source and independently verify key information where appropriate. This content is for informational purposes only and does not constitute legal, financial, investment, or other professional advice.
Sign up for:
Technology, Science, & Me!
The daily local news briefing you can trust. Every day. Subscribe now.
Check Your Email!
We sent a one-time activation link to: .
Confirm it's you by clicking the email link.
If the email is not in your inbox, check spam or try again.
Welcome back!
is already signed up. Check your inbox for updates.