Skip to content
Navigation menu
Search
Powered by Algolia
Search
Log in
Create account
DEV Community
Close
#
evaluation
Follow
Hide
Posts
Left menu
đź‘‹
Sign in
for the ability to sort posts by
relevant
,
latest
, or
top
.
Right menu
An LLM reviewer's "block" is a feature, not a verdict
Cole Halton
Cole Halton
Cole Halton
Follow
Sep 19
An LLM reviewer's "block" is a feature, not a verdict
#
aicodereview
#
llm
#
evaluation
#
machinelearning
Comments
1
 comment
4 min read
My local 7B thinks "kill a Python process" is a violent crime — and my regex beat it
Amirul Cyber
Amirul Cyber
Amirul Cyber
Follow
Sep 17
My local 7B thinks "kill a Python process" is a violent crime — and my regex beat it
#
ai
#
security
#
evaluation
#
llm
Comments
Add Comment
4 min read
RAG Evaluation 2026: The Four Core Metrics and How to Read Them Diagnostically
saaro
saaro
saaro
Follow
Sep 17
RAG Evaluation 2026: The Four Core Metrics and How to Read Them Diagnostically
#
rag
#
ai
#
evaluation
#
metrics
Comments
Add Comment
4 min read
Using Execution Traces to Evaluate AI Agent Behavior
Quantiles.io
Quantiles.io
Quantiles.io
Follow
Sep 17
Using Execution Traces to Evaluate AI Agent Behavior
#
ai
#
opensource
#
evaluation
2
 reactions
Comments
Add Comment
5 min read
Cheap LLM code review is fine until it hits an authorization bug
Cole Halton
Cole Halton
Cole Halton
Follow
Sep 15
Cheap LLM code review is fine until it hits an authorization bug
#
aicodereview
#
llm
#
security
#
evaluation
Comments
Add Comment
2 min read
El modelo encontró evidencia relevante y aun asà falló: por qué “alucinación” se me quedó corta
Cristian Gormaz
Cristian Gormaz
Cristian Gormaz
Follow
Sep 7
El modelo encontró evidencia relevante y aun asà falló: por qué “alucinación” se me quedó corta
#
ai
#
testing
#
llm
#
evaluation
Comments
Add Comment
4 min read
The sleep loop is the tell: agents that pay per action optimize to do nothing
Cole Halton
Cole Halton
Cole Halton
Follow
Sep 7
The sleep loop is the tell: agents that pay per action optimize to do nothing
#
aiagents
#
evaluation
#
llm
#
benchmarking
Comments
1
 comment
2 min read
The model did the reverse-engineering. The validator was the hard part.
Cole Halton
Cole Halton
Cole Halton
Follow
Sep 7
The model did the reverse-engineering. The validator was the hard part.
#
aicoding
#
agents
#
llm
#
evaluation
Comments
1
 comment
2 min read
Judging AI hackathon projects: what to check when every team says 'we used AI'
PRANJUL RATHOUR
PRANJUL RATHOUR
PRANJUL RATHOUR
Follow
Sep 6
Judging AI hackathon projects: what to check when every team says 'we used AI'
#
hackathonjudging
#
ai
#
evaluation
#
rubric
Comments
Add Comment
3 min read
How to evaluate a RAG system: recall, faithfulness and the questions that matter
PRANJUL RATHOUR
PRANJUL RATHOUR
PRANJUL RATHOUR
Follow
Sep 6
How to evaluate a RAG system: recall, faithfulness and the questions that matter
#
rag
#
evaluation
#
metrics
#
guide
Comments
Add Comment
3 min read
The benchmark that disproved its own result
the kilted dev
the kilted dev
the kilted dev
Follow
Sep 10
The benchmark that disproved its own result
#
localmodels
#
llm
#
evaluation
#
buildinpublic
Comments
1
 comment
6 min read
How to Evaluate AI Agents
Quantiles.io
Quantiles.io
Quantiles.io
Follow
Sep 4
How to Evaluate AI Agents
#
ai
#
agentskills
#
agents
#
evaluation
2
 reactions
Comments
Add Comment
7 min read
JuryTrace: make agent-judge failures inspectable
Ama Senevirathne
Ama Senevirathne
Ama Senevirathne
Follow
Sep 5
JuryTrace: make agent-judge failures inspectable
#
ai
#
llm
#
evaluation
#
python
Comments
1
 comment
4 min read
My board never scored an outage as a regression. My evidence couldn't prove it.
Erik Hill
Erik Hill
Erik Hill
Follow
Aug 31
My board never scored an outage as a regression. My evidence couldn't prove it.
#
evaluation
#
testing
#
llm
#
opensource
Comments
Add Comment
5 min read
We spent two days bisecting a prompt change. The regression was noise.
Muhammad Waqas
Muhammad Waqas
Muhammad Waqas
Follow
Aug 30
We spent two days bisecting a prompt change. The regression was noise.
#
mlops
#
regression
#
evaluation
#
agents
Comments
Add Comment
1 min read
đź‘‹
Sign in
for the ability to sort posts by
relevant
,
latest
, or
top
.
We're a place where coders share, stay up-to-date and grow their careers.
Log in
Create account