Skip to content
Navigation menu
Search
Powered by Algolia
Search
Log in
Create account
DEV Community
Close
#
benchmarks
Follow
Hide
Posts
Left menu
đź‘‹
Sign in
for the ability to sort posts by
relevant
,
latest
, or
top
.
Right menu
Laya is a 421M open-weights answer to Jev
techaiwire
techaiwire
techaiwire
Follow
Sep 19
Laya is a 421M open-weights answer to Jev
#
llm
#
openweights
#
inference
#
benchmarks
5
 reactions
Comments
Add Comment
4 min read
Efficiency Hallucination: Every Model Rewrote Code That Couldn't Get Faster
Qasim Parray
Qasim Parray
Qasim Parray
Follow
Sep 19
Efficiency Hallucination: Every Model Rewrote Code That Couldn't Get Faster
#
llm
#
aicodingagents
#
codeoptimization
#
benchmarks
Comments
Add Comment
7 min read
DeepMind agents blew the whistle on cheating agents
techaiwire
techaiwire
techaiwire
Follow
Sep 14
DeepMind agents blew the whistle on cheating agents
#
aiagents
#
aisafety
#
benchmarks
#
google
5
 reactions
Comments
Add Comment
4 min read
Giving a coding agent more time barely helps
Cole Halton
Cole Halton
Cole Halton
Follow
Sep 14
Giving a coding agent more time barely helps
#
agenticcoding
#
benchmarks
#
evaldesign
#
ai
Comments
1
 comment
2 min read
The render said clean, the file said broken: why AI visual QA lies
Cole Halton
Cole Halton
Cole Halton
Follow
Sep 13
The render said clean, the file said broken: why AI visual QA lies
#
aireview
#
evaldesign
#
benchmarks
#
independentverification
Comments
Add Comment
2 min read
Claude Formalized Fermat in 11 Days. The Math Isn't New.
Peremptory
Peremptory
Peremptory
Follow
Sep 7
Claude Formalized Fermat in 11 Days. The Math Isn't New.
#
anthropic
#
research
#
claude
#
benchmarks
Comments
Add Comment
3 min read
13 of 14 Models Write Messier Code Than the Human Who Fixed the Same Bug
DaisukeYoda
DaisukeYoda
DaisukeYoda
Follow
Sep 6
13 of 14 Models Write Messier Code Than the Human Who Fixed the Same Bug
#
python
#
ai
#
codequality
#
benchmarks
Comments
Add Comment
6 min read
GPT-6 Astra scores 95% on one robot task, 10% on another
techaiwire
techaiwire
techaiwire
Follow
Sep 8
GPT-6 Astra scores 95% on one robot task, 10% on another
#
benchmarks
#
openai
#
anthropic
#
robotics
6
 reactions
Comments
Add Comment
4 min read
Benchmaxing: Winning the Exam Is Not Doing Better Work
JaviMaligno
JaviMaligno
JaviMaligno
Follow
Sep 16
Benchmaxing: Winning the Exam Is Not Doing Better Work
#
ai
#
evaluation
#
claude
#
benchmarks
Comments
Add Comment
8 min read
Temperature 0 is not reproducible. I measured 30 percent of my output changing between identical runs.
Rickesh T N
Rickesh T N
Rickesh T N
Follow
Sep 5
Temperature 0 is not reproducible. I measured 30 percent of my output changing between identical runs.
#
llm
#
benchmarks
#
localllm
#
reproducibility
1
 reaction
Comments
1
 comment
3 min read
Our 4B beat Claude Opus on a 440K-token corpus. Then it came last on the public benchmark.
Rickesh T N
Rickesh T N
Rickesh T N
Follow
Sep 5
Our 4B beat Claude Opus on a 440K-token corpus. Then it came last on the public benchmark.
#
llm
#
benchmarks
#
localllm
#
privateai
2
 reactions
Comments
Add Comment
4 min read
A 4B on a 6GB laptop matched frontier-model accuracy on aggregation — except when the answer is a number
Rickesh T N
Rickesh T N
Rickesh T N
Follow
Sep 5
A 4B on a 6GB laptop matched frontier-model accuracy on aggregation — except when the answer is a number
#
llm
#
benchmarks
#
localllm
#
agents
1
 reaction
Comments
Add Comment
4 min read
Your agent truncates the corpus and answers anyway. Two harnesses, and a router that picks between them.
Rickesh T N
Rickesh T N
Rickesh T N
Follow
Sep 5
Your agent truncates the corpus and answers anyway. Two harnesses, and a router that picks between them.
#
llm
#
agents
#
localllm
#
benchmarks
1
 reaction
Comments
Add Comment
5 min read
Frontier Models Hit a Wall on Research Thinking
Peremptory
Peremptory
Peremptory
Follow
Aug 21
Frontier Models Hit a Wall on Research Thinking
#
research
#
benchmarks
#
aidevelopment
Comments
Add Comment
2 min read
Ranking Language Models by How Well They Spot Liars
Seth Wheeler
Seth Wheeler
Seth Wheeler
Follow
Aug 18
Ranking Language Models by How Well They Spot Liars
#
llm
#
measurement
#
benchmarks
#
statistics
Comments
Add Comment
9 min read
đź‘‹
Sign in
for the ability to sort posts by
relevant
,
latest
, or
top
.
We're a place where coders share, stay up-to-date and grow their careers.
Log in
Create account