Psychological methods reveal major weaknesses in AI security testing
Researchers from the UK AI Security Institute have employed psychometric methods to critically assess the efficacy of existing safety benchmarks for language models. Their findings indicate that these benchmarks fail...