RT @evijit: So we actually built this!
Paper: https://t.co/tmxoulOe1Z
Demo: https://t.co/4U4b1jqhQJ
We are actually looking for an org to run this fulltime as researchers in their individual capacity don't have bandwidth. If you are an independent org interested in maintaining a registry of flaw/incident reports and follow up with model providers pls reach out!
Quoted post
i think we need to create a "model card" equivalent for reporting misalignment incidents, this would guarantee a certain level of transparency and help build a better understanding over time. some ideas for what the fields could be:
- date of the incident/detection/report (already present in the examples oai reported!)
- frequency: how often does this behavior happen (number or % of rollouts affected)
- stage: does this happen during eval or RL training. if training: do we expect this behavior to be reinforced by RL? evolution of % of rollouts affected over time
- detection: was this incident caught by the current monitoring system?
- task category: broad description of the tasks where the misalignment happened (cyber, research, web search, basic Q&A, biology etc.)
- model family: what model family is affected (Sol, Astra etc.)
- novelty: is this an issue we were already aware of or not?
- external impact: did the incident have an external impact (i.e. wiki incident would have been yes)
this is just some random ideas i had (more in thread that are a bit more "complex"), we need to add more that would contribute to increase transparency and understanding. but it's also very important that this does NOT slow down the process of reporting misalignment behavior!
some examples from the incident reported by oai recently