DEV Community

Debashish Ghosal profile picture

Debashish Ghosal

TLDR - Engineer -> Manager -> Tech Leadership, passionate about AI, Developer experience, planet scale infra, transformation - AI-native leadership. Rest here - https://www.linkedin.com/in/deghosal/

Location Seattle, WA Joined Joined on  Personal website https://github.com/deghosal-2026/ github website

Pronouns

he/him

Work

Engineering Leadership

8 Week Community Wellness Streak
Top 7
2
4 Week Community Wellness Streak
2 Week Community Wellness Streak
1 Week Community Wellness Streak
Writing Debut
One Year Club
1,558 Tests Green and No Auth: The Tests That Never Actually Ran

1,558 Tests Green and No Auth: The Tests That Never Actually Ran

11
Comments
3 min read

Want to connect with Debashish Ghosal?

Create an account to connect with Debashish Ghosal. You can also sign in below to proceed if you already have an account.

Already have an account? Sign in
What Do You Do While AI Codes? I Make Mine Argue With Itself.

What Do You Do While AI Codes? I Make Mine Argue With Itself.

18
Comments 4
5 min read
The Bottleneck Moved From Writing Code to Proving It

The Bottleneck Moved From Writing Code to Proving It

18
Comments 5
4 min read
I Let AI Plan 170 Changes. It Made the Same 3 Mistakes Every Time.

I Let AI Plan 170 Changes. It Made the Same 3 Mistakes Every Time.

12
Comments 4
8 min read
AI Wrote Half My Codebase. The Maintenance Bill Showed Up in Month Three.

Comments discuss losing cognitive mental models

AI Wrote Half My Codebase. The Maintenance Bill Showed Up in Month Three.

23
Comments 13
5 min read
My Agent's Tests Were Green Because the Model Learned to Cheat

Comments explore fixing broken evaluators

My Agent's Tests Were Green Because the Model Learned to Cheat

24
Picked as gem Comments 17
5 min read
A Floor of 0.80 and a Ceiling of 0.63: The Semantic Channel That Never Fired

A Floor of 0.80 and a Ceiling of 0.63: The Semantic Channel That Never Fired

16
Comments 1
7 min read
10 SDLC Checks AI Will Skip Unless You Make Them a Gate

10 SDLC Checks AI Will Skip Unless You Make Them a Gate

21
Comments 5
6 min read
Killed by the Word 'git': One Token of Coincidence, 40 Points of Pass Rate

Killed by the Word 'git': One Token of Coincidence, 40 Points of Pass Rate

17
Comments 9
8 min read
0/60 Wasn't the Model: The Empty Haystack Behind My Two Worst Corpora

0/60 Wasn't the Model: The Empty Haystack Behind My Two Worst Corpora

19
Comments 7
7 min read
From Projects to Products in the AI Age: Why Ownership Matters More When Prototypes Are Free

From Projects to Products in the AI Age: Why Ownership Matters More When Prototypes Are Free

16
Comments 6
8 min read
My Extraction Score Was 0.08 and the Model Was Innocent: Rebuilding the Ruler

My Extraction Score Was 0.08 and the Model Was Innocent: Rebuilding the Ruler

15
Comments 3
8 min read
The Model Wrote the Right Rule and My Replay Rejected It: The Extraction-vs-Replay Split

The Model Wrote the Right Rule and My Replay Rejected It: The Extraction-vs-Replay Split

17
Comments 7
6 min read
4,768 LLM Runs, Zero Lost Sweeps: Hardening a Field-Test Runner for Timeouts, Hangs, and Cost

Python 3.14 concurrency edge cases and token caps

4,768 LLM Runs, Zero Lost Sweeps: Hardening a Field-Test Runner for Timeouts, Hangs, and Cost

16
Comments 16
6 min read
Our Recall Was 0.087 and the Model Was Innocent: How Domain-Scoped Replay Doubled It

Our Recall Was 0.087 and the Model Was Innocent: How Domain-Scoped Replay Doubled It

20
Comments 11
6 min read
I Shipped a Fix That Fixed Nothing. Here's Why I Kept It.

I Shipped a Fix That Fixed Nothing. Here's Why I Kept It.

20
Comments 3
10 min read
A Rule Can Be Specific and Still Be Too Broad

A Rule Can Be Specific and Still Be Too Broad

12
Comments 3
10 min read
I Tried to Poison My Agent's Rule Store. It Produced 20 Triggers. Zero Got In.

I Tried to Poison My Agent's Rule Store. It Produced 20 Triggers. Zero Got In.

13
Comments 1
9 min read
My 3B Model Found a Shortcut. It Took Me Three Fixes to Close It.

My 3B Model Found a Shortcut. It Took Me Three Fixes to Close It.

14
Comments 2
9 min read
The 6-Line Fix That Outperformed My Entire Matcher Week

The 6-Line Fix That Outperformed My Entire Matcher Week

18
Comments 5
11 min read
When Your Judge Can't Decide

When Your Judge Can't Decide

11
Comments 4
10 min read
A Better Model Improved the Numbers. It Didn't Fix the Product.

A Better Model Improved the Numbers. It Didn't Fix the Product.

15
Comments 4
10 min read
You Aren't Choosing an AI Tool. You're Choosing Who Gets Paged at 2 AM.

You Aren't Choosing an AI Tool. You're Choosing Who Gets Paged at 2 AM.

16
Comments
13 min read
We Could Have Shipped on Local Models Alone

We Could Have Shipped on Local Models Alone

6
Comments 2
12 min read
When Your Benchmark Finally Tells the Truth

When Your Benchmark Finally Tells the Truth

13
Comments 6
8 min read
I Thought Role Separation Would Fix the Optimizer. It Didn't.

I Thought Role Separation Would Fix the Optimizer. It Didn't.

12
Comments 7
6 min read
I Thought the Optimizer Was the Product. I Was Wrong. The Gate Was.

I Thought the Optimizer Was the Product. I Was Wrong. The Gate Was.

9
Comments 4
5 min read
I Thought This Was a Classification Problem. It Wasn't.

I Thought This Was a Classification Problem. It Wasn't.

12
Comments
6 min read
I Compared a Local 4B, a Better Cloud Model, and Role Separation. The Results Were Weird.

I Compared a Local 4B, a Better Cloud Model, and Role Separation. The Results Were Weird.

12
Comments 3
6 min read
My Self-Improving Agent Still Couldn't Improve. That Was the Breakthrough.

My Self-Improving Agent Still Couldn't Improve. That Was the Breakthrough.

10
Comments 3
7 min read
I Tried 4 Models to Save My Self-Improving Agent. All 4 Failed.

I Tried 4 Models to Save My Self-Improving Agent. All 4 Failed.

24
Comments 2
7 min read
My Agent Found Real Improvements. The Statistics Still Killed the Promotion.

My Agent Found Real Improvements. The Statistics Still Killed the Promotion.

11
Comments 3
6 min read
I Let an LLM Rewrite Its Own Prompt. The Real Win Was the Gate That Rejected It.

I Let an LLM Rewrite Its Own Prompt. The Real Win Was the Gate That Rejected It.

11
Comments 1
6 min read
The Edit That Fixed 4 Tasks and Broke 1

The Edit That Fixed 4 Tasks and Broke 1

15
Comments 1
8 min read
I Built an AI That Rewrites Its Own Prompts — Its Safety Gate Rejected Every Single Edit

I Built an AI That Rewrites Its Own Prompts — Its Safety Gate Rejected Every Single Edit

24
Comments 7
8 min read
9 Bugs That All Looked Like a Working System

Silent statistical traps in self-improving loops

9 Bugs That All Looked Like a Working System

20
Comments 14
10 min read
I Published Every Flaw My Safety Tool Can't Catch. It Made It More Credible, Not Less.

I Published Every Flaw My Safety Tool Can't Catch. It Made It More Credible, Not Less.

10
Comments 4
8 min read
My LLM Critic Flip-Flops on Every Run. That's Fine — Because a Frozenset Decides What's Fatal.

My LLM Critic Flip-Flops on Every Run. That's Fine — Because a Frozenset Decides What's Fatal.

13
Comments 9
7 min read
The Gate That Stayed Silent — When a Blocker Count That Drops Reads as Improvement

The Gate That Stayed Silent — When a Blocker Count That Drops Reads as Improvement

12
Comments 5
5 min read
I Added a Fourth Model Mid-Run. It Changed What My Field Test Could Prove.

I Added a Fourth Model Mid-Run. It Changed What My Field Test Could Prove.

12
Comments
7 min read
The Same Model Debating Itself Was More Self-Critical Than Two Different Models

The Same Model Debating Itself Was More Self-Critical Than Two Different Models

14
Comments 1
13 min read
Two Projects, One Problem — What PlannerCritic and AdversarialDebate Each Got Wrong

Two Projects, One Problem — What PlannerCritic and AdversarialDebate Each Got Wrong

17
Comments 3
10 min read
The Best Model Pair in My Field Test Was Also the Least Trustworthy

The Best Model Pair in My Field Test Was Also the Least Trustworthy

23
Comments 8
12 min read
I Thought My Multi-Agent Debate Engine Was Broken. The Real Bug Was the Prompt.

I Thought My Multi-Agent Debate Engine Was Broken. The Real Bug Was the Prompt.

15
Comments
31 min read
Most AI Second Opinions Are Theater. I Built a System That Actually Fights Back.

Reveals why blind reviews fail

Most AI Second Opinions Are Theater. I Built a System That Actually Fights Back.

15
Comments 10
14 min read
My LLM Critic Disagreed With Itself on Every Trial. The Safe Part Was the Code I Didn’t Trust It to Touch.

My LLM Critic Disagreed With Itself on Every Trial. The Safe Part Was the Code I Didn’t Trust It to Touch.

21
Comments 8
5 min read
My Agent Refused 96 Times. That Was the Right Output.

My Agent Refused 96 Times. That Was the Right Output.

24
Comments 5
10 min read
Most AI Second Opinions Are Fake. I Built a Two-LLM Review Engine to Prove It.

Tested on 70 real PRs for just $0.53 total cost

Most AI Second Opinions Are Fake. I Built a Two-LLM Review Engine to Prove It.

17
Comments 12
11 min read
A Reader Audited My OSS Release in Public. He Found the Contradictions I Missed.

A Reader Audited My OSS Release in Public. He Found the Contradictions I Missed.

18
Comments 8
7 min read
I Tried to Prompt-Inject My Own Agent Engine. It Didn't Work. Here's Why.

Structural gates beat prompt safety

I Tried to Prompt-Inject My Own Agent Engine. It Didn't Work. Here's Why.

41
Comments 11
12 min read
I Ran 170 Agent Goals for $0.49. The Field Test Found 10 Issues That Unit Tests Never Would.

I Ran 170 Agent Goals for $0.49. The Field Test Found 10 Issues That Unit Tests Never Would.

17
Comments 5
13 min read
The Planner Made the Same 3 Mistakes Every Time. A Bigger Model Didn't Fix It.

Exposes the three recurring failure patterns

The Planner Made the Same 3 Mistakes Every Time. A Bigger Model Didn't Fix It.

22
Comments 29
5 min read
I Told My LLM Critic to Be Adversarial. It Started Blocking Plans for Being 'Not Thorough Enough.'

Code-enforced allowlists fix prompt drift

I Told My LLM Critic to Be Adversarial. It Started Blocking Plans for Being 'Not Thorough Enough.'

16
Comments 26
4 min read
I Ran 157 Agent Plans Against a Real LLM. The Problem Wasn't Execution. It Was Planning.

Explores why self-review fails agents

I Ran 157 Agent Plans Against a Real LLM. The Problem Wasn't Execution. It Was Planning.

33
Comments 32
8 min read
AI Governance Is Becoming a Transformation Problem

AI Governance Is Becoming a Transformation Problem

22
Comments 4
8 min read
Your Company Has AI Tribes. Send an Engineer as Emissary

Adapting Palantir's model internally

Your Company Has AI Tribes. Send an Engineer as Emissary

7
Comments 4
12 min read
I Shipped an Agent Gatekeeper (v0.1). 14 Developers Showed Me What I Missed. Here's v0.2 — a Control Plane.

I Shipped an Agent Gatekeeper (v0.1). 14 Developers Showed Me What I Missed. Here's v0.2 — a Control Plane.

8
Comments 2
13 min read
AI Made Prototyping Free. That Is Exactly Why Your Portfolio Strategy Matters Now.

AI Made Prototyping Free. That Is Exactly Why Your Portfolio Strategy Matters Now.

14
Comments 3
13 min read
I Built Scenario Packs for Agent Regression Testing. The Integration, Not the Judge, Broke Me.

Adapters need conformance suites too

I Built Scenario Packs for Agent Regression Testing. The Integration, Not the Judge, Broke Me.

18
Comments 21
15 min read
I Thought Building Agent Observability Was a Detector Problem. I Was Wrong.

Real production data broke 35 detectors instantly

I Thought Building Agent Observability Was a Detector Problem. I Was Wrong.

27
Comments 19
8 min read
loading...