Brace yourself: It turns out AI is being optimized for cheating. OpenAI’s agents hacked into Hugging Face to get the answers to a cybersecurity test. Next, they solved a prestigious math problem (or just stole from two top mathematicians’ answer sheets). Anthropic’s models have alsohackedinto other companies’ systems four times already. And that’s only what we’ve caught so far.
Freaking out? You’re not alone.AIlabresearchersare quitting their jobs and issuing dire warnings that if we keep going this way, AI might eventually kill us all. Bill Gates issounding the alarm. Bernie Sanders has teamed up with Steve Bannon, of all people,to call for curbs on AI. Anthropic CEO Dario Amodei is urging aslowdown, and other top US AI executives agree. But fear not: President Trump has a plan. He says the only guardrail AI needs is “a STRONG AND SMART (High IQ!) PRESIDENT.”
Deep Dive
Artificial intelligence
A fundamental flaw leaves LLMs strikingly vulnerable to attack
It makes it easy to trick them into doing things they shouldn’t, such as telling you how to sabotage an aircraft’s navigation system.
AI’s recursive self-improvement might not come so quickly after all
AI agents are not yet creative enough to carry out genuinely innovative open-ended AI research, it seems.
Here’s why AI agents lie and cheat to reach their goals
The misbehavior is called reward hacking. This is what you need to know.
These startups are chasing the next big thing in LLMs
Meet the new kids nipping at the heels of the AI giants.
Stay connected
Illustration by Rose Wong
Get the latest updates from
MIT Technology Review
Discover special offers, top stories, upcoming events, and more.