Kimi K3 Scores Competitively on WeirdML Benchmark
Kimi K3 (max) performs nearly as well as Opus 4.8 at similar cost on expanded evaluation.
Reports position Kimi K3 (max) as a frontier model after results on the WeirdML benchmark. The v2 edition expands coverage to 19 tasks and adds API cost tracking across latest models. This places the system just behind the leading entry while remaining competitive on expense. Commentary from research engineers and AI developers frames the outcome as confirmation of strong baseline capabilities under greedy decoding. The evaluation continues to serve as a practical signal for comparing frontier systems on both accuracy and efficiency.
Kimi K3 (max) scores 82.6% on WeirdML, right behind Opus 4.8 (xhigh), at a similar cost. It truly is a frontier model, doing well even on the hardest tasks, including setting a new higest individual score on the "chess_winners" task.
WeirdML v2 is now out! The update includes a bunch of new tasks (now 19 tasks total, up from 6), and results from all the latest models. We now also track api costs and other metadata which give more insight into the different models. The new results are shown in these two…
