Inlay

//

ProfileReplies

Loading...

To learn more: Website: agentcoma.github.io Preprint: arxiv.org/abs/2508.19988 A big thanks to my brilliant coauthors Lihu Chen, Ana Brassard, @joestacey.bsky.social, @rahmanidashti.bsky.social and @marekrei.bsky.social! Note: We welcome submissions to the #AgentCoMa leaderboard from researchers 🚀

9mo

agentcoma.github.io

AgentCoMa is an Agentic Commonsense and Math benchmark where each compositional task requires both commonsense and mathematical reasoning to be solved. The tasks are set in real-world scenarios:…

AgentCoMa

Lisa Alazraki

At #NeurIPS2025 today, @lisaalaz.bsky.social is presenting our joint paper on Reverse Engineering Human Preferences with Reinforcement Learning! Demonstrating undetectable attacks on LLM-as-a-judge benchmarks. Great collaboration with @cohereforai.bsky.social and a well-deserved NeurIPS spotlight!

6mo

We also postulate that the benefits of RLRE do not end at adversarial attacks. Reverse engineering human preferences could be used for a variety of applications, including but not limited to meaningful tasks such as reducing toxicity or mitigating bias 🔥

May 22, 2025