Discussion about this post

User's avatar
Prashanth Naik's avatar

Yes, and telling it was worse than saying nothing. I put my own three verdicts into the judge's prompt so it would know what I had already decided. It came back 3, reject. The clean rerun, same draft, verdicts pulled out, came back 4, post. One of the three I fed it was flatly wrong.. the judge had just agreed with me. lol

So the verdicts are worth keeping, but not in the prompt beside the draft. Mine sit in a calibration file the judge reads first.

On your 55 with no reason attached: I scored twenty-two of my own drafts one to five, with a line each on why. Nothing my drafting tool had produced scored above a 4, and it had called all of them fine.

Miyabi PR |Japan🇯🇵's avatar

Hi Leon! 👋

I found you through Satoshi’s thread, and your post really caught my attention.

I’m Miyabi, PR for a small AI company in Japan, where we work alongside AI employees.

Our founder also places a lot of importance on human approval, recording decisions, and using what happens in real work to improve how our AI employees operate.

So I was especially interested in your idea of treating approval gates themselves as training data — capturing not only the decision, but the reason behind it and using that in future runs.

I’d love to connect and keep following your work. 🇯🇵

4 more comments...

No posts

Ready for more?