DEV Community

#evaluation

Posts

đź‘‹ Sign in for the ability to sort posts by relevant, latest, or top.
An LLM reviewer's "block" is a feature, not a verdict

An LLM reviewer's "block" is a feature, not a verdict

Comments 1
4 min read
My local 7B thinks "kill a Python process" is a violent crime — and my regex beat it

My local 7B thinks "kill a Python process" is a violent crime — and my regex beat it

Comments
4 min read
RAG Evaluation 2026: The Four Core Metrics and How to Read Them Diagnostically

RAG Evaluation 2026: The Four Core Metrics and How to Read Them Diagnostically

Comments
4 min read
Using Execution Traces to Evaluate AI Agent Behavior

Using Execution Traces to Evaluate AI Agent Behavior

2
Comments
5 min read
Cheap LLM code review is fine until it hits an authorization bug

Cheap LLM code review is fine until it hits an authorization bug

Comments
2 min read
El modelo encontró evidencia relevante y aun así falló: por qué “alucinación” se me quedó corta

El modelo encontró evidencia relevante y aun así falló: por qué “alucinación” se me quedó corta

Comments
4 min read
The sleep loop is the tell: agents that pay per action optimize to do nothing

The sleep loop is the tell: agents that pay per action optimize to do nothing

Comments 1
2 min read
The model did the reverse-engineering. The validator was the hard part.

The model did the reverse-engineering. The validator was the hard part.

Comments 1
2 min read
Judging AI hackathon projects: what to check when every team says 'we used AI'

Judging AI hackathon projects: what to check when every team says 'we used AI'

Comments
3 min read
How to evaluate a RAG system: recall, faithfulness and the questions that matter

How to evaluate a RAG system: recall, faithfulness and the questions that matter

Comments
3 min read
The benchmark that disproved its own result

The benchmark that disproved its own result

Comments 1
6 min read
How to Evaluate AI Agents

How to Evaluate AI Agents

2
Comments
7 min read
JuryTrace: make agent-judge failures inspectable

JuryTrace: make agent-judge failures inspectable

Comments 1
4 min read
My board never scored an outage as a regression. My evidence couldn't prove it.

My board never scored an outage as a regression. My evidence couldn't prove it.

Comments
5 min read
We spent two days bisecting a prompt change. The regression was noise.

We spent two days bisecting a prompt change. The regression was noise.

Comments
1 min read
đź‘‹ Sign in for the ability to sort posts by relevant, latest, or top.