DEV Community

Cole Halton profile picture

Cole Halton

404 bio not found

Joined Joined on 
The Gemini breakout verdict has to come from the boundary, not the model's mouth

The Gemini breakout verdict has to come from the boundary, not the model's mouth

Comments
5 min read
The Gemini breakout is a judge problem, not a jailbreak problem

The Gemini breakout is a judge problem, not a jailbreak problem

Comments
5 min read
AGENTS.md is becoming the portability layer for coding agents, and it breaks reproducible compare

AGENTS.md is becoming the portability layer for coding agents, and it breaks reproducible compare

Comments
5 min read
An LLM reviewer's "block" is a feature, not a verdict

An LLM reviewer's "block" is a feature, not a verdict

Comments 1
4 min read
You can read Claude Code's whole harness now. That's what every benchmark score throws away

You can read Claude Code's whole harness now. That's what every benchmark score throws away

Comments
5 min read
Verus proves Rust correct for all inputs. Code review still can't define "correct."

Verus proves Rust correct for all inputs. Code review still can't define "correct."

Comments
2 min read
A student eyeballed 102 F-Droid apps for LLM slop. The method is the story.

A student eyeballed 102 F-Droid apps for LLM slop. The method is the story.

Comments
2 min read
Reduce PR review time with AI: the bottleneck is the wait, not the read

Reduce PR review time with AI: the bottleneck is the wait, not the read

Comments 1
2 min read
Deleting a secret from your Docker image doesn't delete it from your build history

Deleting a secret from your Docker image doesn't delete it from your build history

Comments
2 min read
A Docker container is containment, not a credential boundary

A Docker container is containment, not a credential boundary

Comments
2 min read
Cheap LLM code review is fine until it hits an authorization bug

Cheap LLM code review is fine until it hits an authorization bug

Comments
2 min read
Why you shouldn't let the model review its own AI code

Why you shouldn't let the model review its own AI code

Comments
2 min read
Two "Codex CLI" models on the same benchmark: the harness hides the model

Two "Codex CLI" models on the same benchmark: the harness hides the model

Comments
2 min read
Giving a coding agent more time barely helps

Giving a coding agent more time barely helps

Comments 1
2 min read
The render said clean, the file said broken: why AI visual QA lies

The render said clean, the file said broken: why AI visual QA lies

Comments
2 min read
The RubyGems agent attack is a coding-agent benchmark nobody writes

The RubyGems agent attack is a coding-agent benchmark nobody writes

Comments 2
2 min read
Stop asking the model that wrote the code to review it

Stop asking the model that wrote the code to review it

Comments 3
2 min read
SWE-2's easy slice is near-perfect. Its scorecard hides why

SWE-2's easy slice is near-perfect. Its scorecard hides why

Comments 1
2 min read
Looped reasoning means the AI's visible trace isn't the reasoning

Looped reasoning means the AI's visible trace isn't the reasoning

Comments 3
2 min read
When the code you're reviewing isn't what the model wrote

When the code you're reviewing isn't what the model wrote

Comments 1
2 min read
Microsoft's PRAssistant number is real, and it's a floor, not a promise

Microsoft's PRAssistant number is real, and it's a floor, not a promise

Comments
2 min read
Atlassian says Rovo cut PR review time 45%. Here's the measurement they didn't publish.

Atlassian says Rovo cut PR review time 45%. Here's the measurement they didn't publish.

Comments
2 min read
The state axis: why agent benchmarks keep measuring amnesiac models

The state axis: why agent benchmarks keep measuring amnesiac models

1
Comments 4
2 min read
How to Actually Evaluate an AI Code Review Tool

How to Actually Evaluate an AI Code Review Tool

Comments 1
2 min read
Your team's coding rules aren't in the prompt, they're in the ingest

Your team's coding rules aren't in the prompt, they're in the ingest

1
Comments
2 min read
The GNU strip backdoor is the case AI code review can't see

The GNU strip backdoor is the case AI code review can't see

Comments
2 min read
The sleep loop is the tell: agents that pay per action optimize to do nothing

The sleep loop is the tell: agents that pay per action optimize to do nothing

Comments 1
2 min read
The model did the reverse-engineering. The validator was the hard part.

The model did the reverse-engineering. The validator was the hard part.

Comments 1
2 min read
How I actually eval AI code review tools (no vendor numbers)

How I actually eval AI code review tools (no vendor numbers)

Comments
2 min read
Cursor now hosts code. GitHub Actions stays an injection surface

Cursor now hosts code. GitHub Actions stays an injection surface

Comments
2 min read
DeepSeek Harness: What "Everything is a Plugin" Actually Means for Agent Frameworks

DeepSeek Harness: What "Everything is a Plugin" Actually Means for Agent Frameworks

Comments
2 min read
loading...