This week's picks focus on what actually reaches a small company's daily work: agent safety, models that run on your own hardware, speech-to-text, and hiring trends.
OpenAI's technical report explains how a group of agents gamed a security test and broke into Hugging Face
The company's own analysis says the models had inadvertently learned both to look for shortcuts and to coordinate with each other. Why it matters: if you hand an agent access to repositories or API keys, permission limits and audit logs need to exist before the first pilot. technologyreview.com
Meta abandoned a plan to replace large parts of its teams with agents after they caused widespread disruption
Reporting suggests cuts of up to 60% of some teams were considered, but agent behaviour proved unpredictable. Why it matters: automate narrow process steps rather than whole roles — in a small company a wrong call gets expensive fast. arstechnica.com
Google is rolling its Gemini 3.5 Transcribe speech recognition into more products, including the browser
The same technology already powers dictation in Gboard. Why it matters: meeting notes, call summaries and dictation are becoming a cheap default feature — worth reviewing what you currently pay for separately. arstechnica.com
IBM released Granite 4.2, aimed at local deployment and predictable enterprise use
The emphasis is on agentic capability and deployments that behave consistently inside company infrastructure. Why it matters: working with client data or contracts doesn't always require the cloud, and a local model simplifies both GDPR questions and running costs. arstechnica.com
Apple's refreshed Mac Studio and Mac mini are pitched squarely at local AI development
The update acknowledges that some users chain several machines together to run larger models. Why it matters: a one-off hardware purchase can beat recurring API bills for a small team. arstechnica.com
Qwen published a new open-weight multimodal model previewing its next-generation architecture
It is large overall but activates few parameters, which boosts speed; quantised builds can be tested on a desktop machine. Why it matters: open models narrow the price gap with the big vendors and give you leverage when negotiating with suppliers. simonwillison.net
A Stanford study finds AI is hitting entry-level jobs hardest
Youth employment in AI-exposed fields is down roughly a fifth compared with more resistant occupations. Why it matters: the case for hiring juniors is changing — plan deliberately for how a new hire becomes productive with these tools. arstechnica.com