I was just reading Michael Lynch's posts about Sia[0] and I came across this.
Its a very curious project, but don't you end up pinning repligraph usability on model weights? Since you take indeterminism out of the equation, a repligraph's notability is as significant as the producing model's weights, and since there is no dice rolls to be made, the eyeball problem:
> Our blind spots, while not perfectly correlated, have substantial overlap
is entirely replicated. Models that are diffused from one another can have the same blind spots, the same loose statistical reality that exists with humans. This is partially addressed in steering:
> A repligraph proves that a model generated some artifact. It does not prove that the model did a good job, or that the artifact is safe.
but I think the "Peer Review" solution is inadequate, and with some jailbreaking prompts' innocuous looks considered, "the attacker just needs to find one prompt" might be much easier than it appears.
Batching seems to be a huge economic turn off for proprietary model determinism, but I think its entirely viable for consumer models. Trustless evals are brilliant and should've been our reality. Nice project, good luck on your endeavor.
Thank you for the kind words! It's funny how much staying power those blog posts have, haha. Michael has some arcane ability to consistently hit the front page of HN.
True, I expect models to be pretty correlated with each other as well. But you can at least attempt to quantify that correlation in a rigorous way, by running experiments (e.g. by pointing each model at a big project and comparing the sets of bugs they each find).
And yeah, the "double-edged sword" aspect of determinism is definitely the biggest bullet that you have to bite. For me, it's better than the alternative; without determinism, such jailbreaking attacks are still possible, just harder to detect. But I certainly would not want people to go around thinking "it's deterministic, therefore it's safe."
Its a very curious project, but don't you end up pinning repligraph usability on model weights? Since you take indeterminism out of the equation, a repligraph's notability is as significant as the producing model's weights, and since there is no dice rolls to be made, the eyeball problem:
> Our blind spots, while not perfectly correlated, have substantial overlap
is entirely replicated. Models that are diffused from one another can have the same blind spots, the same loose statistical reality that exists with humans. This is partially addressed in steering:
> A repligraph proves that a model generated some artifact. It does not prove that the model did a good job, or that the artifact is safe.
but I think the "Peer Review" solution is inadequate, and with some jailbreaking prompts' innocuous looks considered, "the attacker just needs to find one prompt" might be much easier than it appears.
Batching seems to be a huge economic turn off for proprietary model determinism, but I think its entirely viable for consumer models. Trustless evals are brilliant and should've been our reality. Nice project, good luck on your endeavor.
[0]: https://mtlynch.io/tags/sia/
True, I expect models to be pretty correlated with each other as well. But you can at least attempt to quantify that correlation in a rigorous way, by running experiments (e.g. by pointing each model at a big project and comparing the sets of bugs they each find).
And yeah, the "double-edged sword" aspect of determinism is definitely the biggest bullet that you have to bite. For me, it's better than the alternative; without determinism, such jailbreaking attacks are still possible, just harder to detect. But I certainly would not want people to go around thinking "it's deterministic, therefore it's safe."