DEV Community

Testing

Find those bugs before your users do! 🐛

Posts

👋 Sign in for the ability to sort posts by relevant, latest, or top.
389 Tests Passed. NIST Still Caught the Bug.

NIST validation exposes limits of unit tests

389 Tests Passed. NIST Still Caught the Bug.

4
Comments 6
7 min read
I built a tool to prove my multi-agent harness was worth it. It told me it wasn't.

I built a tool to prove my multi-agent harness was worth it. It told me it wasn't.

1
Comments 2
4 min read
I Replaced pytest-xdist With 60 Lines of subprocess.Popen. Here's Why.

I Replaced pytest-xdist With 60 Lines of subprocess.Popen. Here's Why.

Comments
4 min read
Stress-testing my Multi-LLM engine: 93 chunks, 8 models, and one "Insufficient Balance" error.

Stress-testing my Multi-LLM engine: 93 chunks, 8 models, and one "Insufficient Balance" error.

Comments
2 min read
When the Mock Is Right and the API Isn't Anymore

When the Mock Is Right and the API Isn't Anymore

Comments
2 min read
Our LLM Judges Called Human Writing "AI-Flavored" 88% of the Time

Our LLM Judges Called Human Writing "AI-Flavored" 88% of the Time

Comments
3 min read
The Test That Only Failed in CI, Never Locally

The Test That Only Failed in CI, Never Locally

Comments
2 min read
From Bug Found to Bug Filed: A Bug-Reporter Skill for Claude Code

From Bug Found to Bug Filed: A Bug-Reporter Skill for Claude Code

Comments
5 min read
Test Result Reporting and Failing Fast in CI Pipelines

Test Result Reporting and Failing Fast in CI Pipelines

Comments
7 min read
Building a Timing Utility That Can't Corrupt Its Own Stats — Even When Your Code Throws

Building a Timing Utility That Can't Corrupt Its Own Stats — Even When Your Code Throws

Comments
4 min read
My marketing docs kept lying, so I put them in CI

My marketing docs kept lying, so I put them in CI

Comments
2 min read
How to Generate Realistic E-Commerce Test Data in 2 Lines of Code

How to Generate Realistic E-Commerce Test Data in 2 Lines of Code

2
Comments
3 min read
Stop letting your coding agent claim "pixel-perfect". Make it prove 97.49%.

Stop letting your coding agent claim "pixel-perfect". Make it prove 97.49%.

Comments 2
2 min read
The Biggest Flaw in My AI Evaluation Wasn't the Models. It Was My Scorecard.

The Biggest Flaw in My AI Evaluation Wasn't the Models. It Was My Scorecard.

Comments
2 min read
Design Around the Point You Cannot Undo

Design Around the Point You Cannot Undo

Comments
4 min read
👋 Sign in for the ability to sort posts by relevant, latest, or top.