back

by walrus01·10d ago·view on hn ↗
It would be an interesting comparison to compare the human "miss rate" shown in the table there with the exact same tests repeated with a different LLM watching and approving or denying each action. No human in the loop, just record the results and take the measurement of pass/fail at the end of the run. With something fairly large and smart that has been given a very specific system prompt to watch and prevent harmful actions or data leaks.