12 votes

Natural language autoencoders

2 comments

  1. pete_the_paper_boat
    (edited )
    Link

    NLAs hallucinate. NLA explanations frequently contain claims about the context that are verifiably false. These are often easy to catch by checking against the transcript, but the same failure could extend to claims about the model's internal processing, which are harder to verify. This makes NLAs difficult to rely on.

    We find that NLA runs do not consistently help the auditor discover more behavioral quirks in the target model. We see a noticeable but small bump in performance according to the Rubric (out of four possible points) for NLA runs.

    2 votes
  2. balooga
    Link
    This kind of research is exciting to me! Anything that demystifies the “black box” — even if imperfectly — is a step in the right direction. It also has the potential to improve the safety and...

    This kind of research is exciting to me! Anything that demystifies the “black box” — even if imperfectly — is a step in the right direction. It also has the potential to improve the safety and alignment of future models.

    On the other hand, probing the thoughts of AIs to suss out their motivations and secrets feels kinda Blade Runner… it’s like the Voight-Kampff test in real life. Strange times we're living in.

    1 vote