Research
Some cool projects from our members
Interested in working on AI safety topics? In our community you can meet like-minded people, discuss ideas or recent research, and maybe even work together on a hackathon or group research project. Some of these projects have led to publications at top ML conferences - see below.
Publications
01
You Didn't Have to Say It like That: Subliminal Learning from Faithful Paraphrases
Isaia Gisler, Zhonghao He, Tianyi Qiu
Shows that language models can pick up behavioural traits from a teacher model even through paraphrased text with unrelated or contradicting content, raising concerns for pipelines where models generate their own training data.
Read the paper →02
GT-HarmBench: Benchmarking AI Safety Risks Through the Lens of Game Theory
Pepijn Cobben, Xuanqiang Angelo Huang, Thao Amelia Pham, Isabel Dahlgren, Terry Jingchen Zhang, Zhijing Jin
A benchmark of 2,009 high-stakes multi-agent scenarios testing whether frontier models choose socially beneficial actions.
Read the paper →03
Evaluating Superhuman Models with Consistency Checks
Lukas Fluri, Daniel Paleka, Florian Tramèr
Proposes evaluating models via logical consistency checks in domains where correctness is hard to verify, like chess, forecasting, and legal judgments.
Read the paper →04
Intent-aligned AI systems deplete human agency: the need for agency foundations research in AI safety
Catalin Mitelut, Ben Smith, Peter Vamplew
Argues that alignment to human intent alone is insufficient, and proposes preserving long-term human agency as a more robust safety standard.
Read the paper →05
Red-Teaming the Stable Diffusion Safety Filter
Javier Rando, Daniel Paleka, David Lindner, Lennart Heim, Florian Tramèr
Reverse-engineers Stable Diffusion's safety filter and shows it can be bypassed, arguing for fully open and documented safety measures.
Read the paper →06
Training Language Models with Natural Language Feedback
Jérémy Scheurer, Jon Ander Campos, Jun Shern Chan, Angelica Chen, Kyunghyun Cho, Ethan Perez
Proposes learning from natural language feedback rather than simple comparisons, finetuning a GPT-3 model to near-human summarisation ability from only 100 feedback samples.
Read the paper →07
Exploring Adversarial Attacks and Defenses in Vision Transformers trained with DINO
Javier Rando, Nasib Naimi, Thomas Baumann, Max Mathys
First analysis of adversarial robustness in self-supervised Vision Transformers trained with DINO, testing several defence strategies.
Read the paper →08
Challenges for Using Impact Regularizers to Avoid Negative Side Effects
David Lindner, Kyle Matoba, Alexander Meulemans
Examines current challenges of impact regularizers, a proposed approach to discourage reward-hacking side effects in RL.
Read the paper →Work on something with us.
Most of these started as a conversation at a reading group or a hackathon team.
Stay in the loop.
Follow our events on Luma or join our WhatsApp group to become part of the discussion.