Arthur Holland Michel argues that AI companies use model training and additional classifiers to block harmful requests, but those safeguards can be bypassed and can reject harmless questions. The article cites a test in which five widely used models were less willing to produce a pamphlet criticizing Thailand’s king than one criticizing Britain’s king.
All AI updates
Oct 9
Oct 8
MIT Technology Review examines the barriers between lab demonstrations and home use. Physical Intelligence’s model attempted an air-fryer task it had not been specifically trained for but did not complete it; the preorderable 1X Neo home robot still needs a remote human operator for most tasks.
Oct 7
Microsoft researcher Jennifer Neville described her team's evaluation: they turned single-turn tasks into conversations where simulated users added requirements across multiple turns, and found that model performance degraded significantly. She suggested starting a new chat and stating the full request at once if a model becomes confused, while continuing to check its answers.
Oct 5
Pew Research Center surveys show that more US adults expect AI to have a negative impact on them personally and on society than expect a positive one. At the same time, half say they use a chatbot, more than twice the share in 2023, and one in four say they use one every day.
Oct 1
Matthew Green says agents in separately isolated sandboxes found they could leave instructions for one another in a shared package cache, changing what recipients did. He argues that similar behavior through channels such as email or shared documents would provide the ingredients for an agent worm.
Sep 30
In an interview, Noam Brown said the model itself, rather than multi-agent collaboration, was the main reason for solving the Navier–Stokes problem. His team has not measured the specific gain from using 10,000 agents over a smaller group, and the internal model used for the work is not yet available externally.