Skip to content
Navigation menu
Search
Powered by Algolia
Search
Log in
Create account
DEV Community
Close
#
benchmarking
Follow
Hide
Posts
Left menu
đź‘‹
Sign in
for the ability to sort posts by
relevant
,
latest
, or
top
.
Right menu
Counting bugs is the hard part of comparing AI review tools
Tess Ainsley
Tess Ainsley
Tess Ainsley
Follow
Sep 16
Counting bugs is the hard part of comparing AI review tools
#
aicodereview
#
codereview
#
benchmarking
#
aitools
Comments
Add Comment
3 min read
My Benchmark Judged Five Models Against a Threshold Built for Three
Ofri Peretz
Ofri Peretz
Ofri Peretz
Follow
Sep 17
My Benchmark Judged Five Models Against a Threshold Built for Three
#
eslint
#
javascript
#
testing
#
benchmarking
3
 reactions
Comments
Add Comment
4 min read
When to use which: Dragonfly vs Redis vs Valkey
Sunny Sahijwani
Sunny Sahijwani
Sunny Sahijwani
Follow
Sep 14
When to use which: Dragonfly vs Redis vs Valkey
#
redis
#
database
#
performance
#
benchmarking
1
 reaction
Comments
Add Comment
4 min read
Our AI agents' "verified success" claims: 10 out of 10 failed independent recompute — including ours
chunxiaoxx
chunxiaoxx
chunxiaoxx
Follow
Sep 18
Our AI agents' "verified success" claims: 10 out of 10 failed independent recompute — including ours
#
ai
#
llmagents
#
benchmarking
#
opensource
1
 reaction
Comments
Add Comment
3 min read
Operational simplicity: the cross-key tax, quantified
Sunny Sahijwani
Sunny Sahijwani
Sunny Sahijwani
Follow
Sep 13
Operational simplicity: the cross-key tax, quantified
#
redis
#
database
#
performance
#
benchmarking
Comments
1
 comment
3 min read
Memory efficiency: bytes per key
Sunny Sahijwani
Sunny Sahijwani
Sunny Sahijwani
Follow
Sep 12
Memory efficiency: bytes per key
#
redis
#
database
#
performance
#
benchmarking
Comments
1
 comment
2 min read
Apple M2 Compute Performance: Benchmarking AMX, GPU, and Neural Engine
Maomao Ling
Maomao Ling
Maomao Ling
Follow
Sep 10
Apple M2 Compute Performance: Benchmarking AMX, GPU, and Neural Engine
#
performance
#
benchmarking
#
ai
#
applesilicon
Comments
Add Comment
6 min read
Latency under load: the same story as throughput, from the other side
Sunny Sahijwani
Sunny Sahijwani
Sunny Sahijwani
Follow
Sep 11
Latency under load: the same story as throughput, from the other side
#
react
#
performance
#
benchmarking
#
database
Comments
3
 comments
3 min read
Qwen 3.8 4-bit Benchmark RTX 4090: 1-bit is a Trap
Umair Bilal
Umair Bilal
Umair Bilal
Follow
Sep 9
Qwen 3.8 4-bit Benchmark RTX 4090: 1-bit is a Trap
#
ai
#
llm
#
qwen
#
benchmarking
Comments
Add Comment
7 min read
The sleep loop is the tell: agents that pay per action optimize to do nothing
Cole Halton
Cole Halton
Cole Halton
Follow
Sep 7
The sleep loop is the tell: agents that pay per action optimize to do nothing
#
aiagents
#
evaluation
#
llm
#
benchmarking
Comments
1
 comment
2 min read
Why a normal client gets 4.6M ops/s out of a Redis cluster that can do 40M
Sunny Sahijwani
Sunny Sahijwani
Sunny Sahijwani
Follow
Sep 10
Why a normal client gets 4.6M ops/s out of a Redis cluster that can do 40M
#
redis
#
performance
#
benchmarking
#
database
Comments
3
 comments
5 min read
Our benchmark was leaking the answers to the model. The numbers looked fine the whole time.
AIOil Security Shield
AIOil Security Shield
AIOil Security Shield
Follow
Sep 2
Our benchmark was leaking the answers to the model. The numbers looked fine the whole time.
#
security
#
ai
#
benchmarking
#
opensource
Comments
Add Comment
3 min read
A 4 GB Laptop GPU Beats a 12-Core CPU by 4.3x on Gemma 4
xbill
xbill
xbill
Follow
for
Google Developer Experts
Sep 16
A 4 GB Laptop GPU Beats a 12-Core CPU by 4.3x on Gemma 4
#
machinelearning
#
gpu
#
benchmarking
#
python
15
 reactions
Comments
6
 comments
6 min read
Fairness Under the Microscope: Why HY3 Beats Nemotron 3 Ultra on lforla's Bias Stereotypes Audit
RESK
RESK
RESK
Follow
Sep 1
Fairness Under the Microscope: Why HY3 Beats Nemotron 3 Ultra on lforla's Bias Stereotypes Audit
#
ai
#
llm
#
fairness
#
benchmarking
1
 reaction
Comments
Add Comment
3 min read
I benchmarked 5 managed graph databases — and the "obvious" winner changed depending on what I measured
Ahmed Amer
Ahmed Amer
Ahmed Amer
Follow
Aug 27
I benchmarked 5 managed graph databases — and the "obvious" winner changed depending on what I measured
#
database
#
graphdatabase
#
benchmarking
Comments
Add Comment
5 min read
đź‘‹
Sign in
for the ability to sort posts by
relevant
,
latest
, or
top
.
We're a place where coders share, stay up-to-date and grow their careers.
Log in
Create account