Paraphernalia
AarXiv2023

Conditioning Predictive Models: Risks and Strategies

Evan Hubinger, Adam S. Jermyn, Johannes Treutlein, Rubi Hudson, Kate Woolverton

Abstract

'Kate Woolverton'] Unfortunately, such approaches also raise a variety of potentially fatal safety problems, particularly surrounding situations where predictive models predict the output of other AI systems, potentially unbeknownst to us. There are numerous potential solutions to such problems, however, primarily via carefully conditioning models to predict the things we want—e.g. humans—rather than the things we don't—e.g. malign AIs.

§ The Valyu brief

Reading the full paper and taking notes. This takes a few seconds…

§ Ask this paper

Ask a question about this paper

Valyu reads the full text and answers from what the paper actually says.

Q.

Searching the other archives…

Conditioning Predictive Models: Risks and Strategies · Paraphernalia