Someone embedded hidden instructions in a court document designed to manipulate any AI system that processed it — essentially hacking the judge's AI assistant. This is the first documented case of prompt injection entering the legal system, and it won't be the last.
Twitch quietly enabled a setting that feeds creator content into Amazon's AI training pipeline. The opt-out exists, but most streamers have no idea it's on.
Simon Willison's LLM plugin now supports Gemini 3.7 Flash, which Google dropped this week — 50% cheaper than 3.6 Flash and meaningfully smarter. The pace of Flash-tier iteration is remarkable: three generations in roughly as many weeks.
The latest DeepSeek Pro model dropped via API with no official announcement page — just a quiet push to OpenRouter. Open weights status remains unconfirmed, which is the only number that matters to the open-source crowd.
Mirendil cofounders — ex-Google and Anthropic researchers — are building self-accelerating AI systems that contribute to their own development. The question of when AI meaningfully speeds up its own progress is moving from philosophy to product roadmap.
Datadog's security chief explains how a company with 4,000+ engineers using coding agents doesn't try to block AI — it builds infrastructure around it. The embrace-and-secure playbook is becoming the enterprise standard.
Google's Gemini product lead announced Gemini 3.7 Flash with a sharp stat: 50% price cut versus 3.6 Flash, shipping roughly three weeks after its predecessor. When the gap between model generations shrinks to weeks, the competitive landscape stops being about who has the best model and starts being about who can ship fastest.
A research decomposition of major agent benchmarks — TheAgentCompany and others — found far less signal in the rankings than the leaderboard positions imply. Translation: the agent benchmark arms race may be measuring noise. Caveat everything you've read about agent performance this year.
Infisical's approach to agentic security is elegant: give the agent a fake API key, then swap in the real credential at execution time. Agents never hold your actual secrets. This kind of zero-trust-for-agents architecture is going to become table stakes fast.
Mollick pushed back on the emerging "returns to AI are plateauing" narrative, arguing that people are systematically underestimating what higher-intelligence models unlock. He's also trolling the AI commentariat for suddenly pretending they always understood the nuances of mathematical benchmarks — a fair hit.
Two stories this morning are easy to underreact to. A prompt injection hidden inside a legal filing is not a curiosity — it's a proof of concept for a new category of adversarial document. Courts are adopting AI-assisted review tools precisely as bad actors learn to weaponize them. The legal system is about to get a crash course in a security threat that the tech industry has known about for two years and mostly ignored.
Meanwhile the agent benchmark trust crisis is real and underreported. If DAIR.AI's decomposition holds up, much of the agent performance comparison that drives vendor decisions and investor theses is built on noisy data. That's a problem when the entire enterprise security conversation — as seen in the Datadog CISO interview — is premised on understanding which agents can actually be trusted with production access. And Twitch quietly flipping on AI training for streamer content without fanfare is a reminder that the data acquisition war never stopped; it just got quieter.