DEV Community

#benchmarks

Posts

đź‘‹ Sign in for the ability to sort posts by relevant, latest, or top.
Laya is a 421M open-weights answer to Jev

Laya is a 421M open-weights answer to Jev

5
Comments
4 min read
Efficiency Hallucination: Every Model Rewrote Code That Couldn't Get Faster

Efficiency Hallucination: Every Model Rewrote Code That Couldn't Get Faster

Comments
7 min read
DeepMind agents blew the whistle on cheating agents

DeepMind agents blew the whistle on cheating agents

5
Comments
4 min read
Giving a coding agent more time barely helps

Giving a coding agent more time barely helps

Comments 1
2 min read
The render said clean, the file said broken: why AI visual QA lies

The render said clean, the file said broken: why AI visual QA lies

Comments
2 min read
Claude Formalized Fermat in 11 Days. The Math Isn't New.

Claude Formalized Fermat in 11 Days. The Math Isn't New.

Comments
3 min read
13 of 14 Models Write Messier Code Than the Human Who Fixed the Same Bug

13 of 14 Models Write Messier Code Than the Human Who Fixed the Same Bug

Comments
6 min read
GPT-6 Astra scores 95% on one robot task, 10% on another

GPT-6 Astra scores 95% on one robot task, 10% on another

6
Comments
4 min read
Benchmaxing: Winning the Exam Is Not Doing Better Work

Benchmaxing: Winning the Exam Is Not Doing Better Work

Comments
8 min read
Temperature 0 is not reproducible. I measured 30 percent of my output changing between identical runs.

Temperature 0 is not reproducible. I measured 30 percent of my output changing between identical runs.

1
Comments 1
3 min read
Our 4B beat Claude Opus on a 440K-token corpus. Then it came last on the public benchmark.

Our 4B beat Claude Opus on a 440K-token corpus. Then it came last on the public benchmark.

2
Comments
4 min read
A 4B on a 6GB laptop matched frontier-model accuracy on aggregation — except when the answer is a number

A 4B on a 6GB laptop matched frontier-model accuracy on aggregation — except when the answer is a number

1
Comments
4 min read
Your agent truncates the corpus and answers anyway. Two harnesses, and a router that picks between them.

Your agent truncates the corpus and answers anyway. Two harnesses, and a router that picks between them.

1
Comments
5 min read
Frontier Models Hit a Wall on Research Thinking

Frontier Models Hit a Wall on Research Thinking

Comments
2 min read
Ranking Language Models by How Well They Spot Liars

Ranking Language Models by How Well They Spot Liars

Comments
9 min read
đź‘‹ Sign in for the ability to sort posts by relevant, latest, or top.