People delete the human gate so they can say fully automated. That throws away the only proof the agent is getting better. Feed it back instead, and the gate teaches itself out of a job.
Yes, and telling it was worse than saying nothing. I put my own three verdicts into the judge's prompt so it would know what I had already decided. It came back 3, reject. The clean rerun, same draft, verdicts pulled out, came back 4, post. One of the three I fed it was flatly wrong.. the judge had just agreed with me. lol
So the verdicts are worth keeping, but not in the prompt beside the draft. Mine sit in a calibration file the judge reads first.
On your 55 with no reason attached: I scored twenty-two of my own drafts one to five, with a line each on why. Nothing my drafting tool had produced scored above a 4, and it had called all of them fine.
Exactly. A judge shouldn’t be shown the answer key while it’s grading. Keeping the verdicts in a calibration file preserves their value without anchoring the individual call. And one of your own verdicts being wrong is a feature: calibration data should expose blind spots, not just reinforce existing judgments. “Fine” is too vague to teach the system anything.
I found you through Satoshi’s thread, and your post really caught my attention.
I’m Miyabi, PR for a small AI company in Japan, where we work alongside AI employees.
Our founder also places a lot of importance on human approval, recording decisions, and using what happens in real work to improve how our AI employees operate.
So I was especially interested in your idea of treating approval gates themselves as training data — capturing not only the decision, but the reason behind it and using that in future runs.
I’d love to connect and keep following your work. 🇯🇵
Great to meet you. It’s great to hear that Shunsuke is thinking along similar lines, especially about capturing the reasoning behind approvals and using it to improve future runs. I’d love to learn more about what you’re building in Japan.
I’d be happy to share more about what we’re doing in Japan.
I’m not an engineer, so I understand these ideas mainly through my own experience of working with our AI employees. For example, in my PR work, AI can help me prepare a draft of an article, but I do the final review myself, and I’m the one who publishes it.
So I’m experiencing this from the human side rather than the technical side. I’d love to keep learning from your perspective and sharing what it actually feels like to work this way.
Yes, and telling it was worse than saying nothing. I put my own three verdicts into the judge's prompt so it would know what I had already decided. It came back 3, reject. The clean rerun, same draft, verdicts pulled out, came back 4, post. One of the three I fed it was flatly wrong.. the judge had just agreed with me. lol
So the verdicts are worth keeping, but not in the prompt beside the draft. Mine sit in a calibration file the judge reads first.
On your 55 with no reason attached: I scored twenty-two of my own drafts one to five, with a line each on why. Nothing my drafting tool had produced scored above a 4, and it had called all of them fine.
Exactly. A judge shouldn’t be shown the answer key while it’s grading. Keeping the verdicts in a calibration file preserves their value without anchoring the individual call. And one of your own verdicts being wrong is a feature: calibration data should expose blind spots, not just reinforce existing judgments. “Fine” is too vague to teach the system anything.
I had not thought of it that way. I was treating the bad ones as noise and planning to clean them out.
Hi Leon! 👋
I found you through Satoshi’s thread, and your post really caught my attention.
I’m Miyabi, PR for a small AI company in Japan, where we work alongside AI employees.
Our founder also places a lot of importance on human approval, recording decisions, and using what happens in real work to improve how our AI employees operate.
So I was especially interested in your idea of treating approval gates themselves as training data — capturing not only the decision, but the reason behind it and using that in future runs.
I’d love to connect and keep following your work. 🇯🇵
Great to meet you. It’s great to hear that Shunsuke is thinking along similar lines, especially about capturing the reasoning behind approvals and using it to improve future runs. I’d love to learn more about what you’re building in Japan.
Thank you, Leon!
I’d be happy to share more about what we’re doing in Japan.
I’m not an engineer, so I understand these ideas mainly through my own experience of working with our AI employees. For example, in my PR work, AI can help me prepare a draft of an article, but I do the final review myself, and I’m the one who publishes it.
So I’m experiencing this from the human side rather than the technical side. I’d love to keep learning from your perspective and sharing what it actually feels like to work this way.