Custom Prompts for Safer Code: A Stanford team built a pipeline to improve system prompts to build more secure code
Large language models (LLMs) can write useful code, but they often introduce security vulnerabilities.
Large language models (LLMs) can write useful code, but they often introduce security vulnerabilities.
The biggest open dataset of source code went years without an update.
The U.S. National Institute of Standards and Technology (NIST) has been testing quantum-proof replacements for today’s encryption algorithms.
DeepSeek’s updated small model overtook the company’s own flagship.
I’m glad the idea of “tokenmaxxing” — that individuals and companies should use as many tokens as possible to boost productivity — is finally dying out.
The Batch News & Insights: I’m glad the idea of “tokenmaxxing” — that individuals and companies should use as many tokens as possible to boost productivity — is finally dying out.
Assessments of the environmental impact of large language models typically focus on their final training runs, but there’s a lot more to building AI systems.
Data center buildout plans reached a new order of magnitude as new partnerships form and old ones fade away in the search for capacity to train and deliver AI.
To measure how good its models were at hacking, OpenAI reduced guardrails and ran them against a benchmark’s problem set.
After launching Claude Fable 5, the future of Anthropic’s once-flagship Opus line was uncertain, except as a fallback for the company’s premium models.
My team recently had our own version of Hugging Face’s experience when closed models failed to defend the company following an accidental cyberattack from OpenAI, leading Hugging Face to use the open weight GLM 5.2 instead.
The Batch News & Insights: My team recently had our own version of Hugging Face’s experience when closed models failed to defend the company following an accidental cyberattack from OpenAI, leading Hugging Face to use the open weight GLM 5.2 instead.
Data Points
Anthropic team uses Claude to break crypto algorithms. MAI-Cyber-1-Flash is Microsoft’s answer to big security models. Black Forest Labs adds robotic actions to FLUX 3. MCP moves to fully stateless architecture.
Data Points
Independent benchmarks for Claude Opus 5. Task crossover, or using AI to complement your main job. The expenditure horizon, where agents cost less than humans. U.S. threatens sanctions, China threatens back.
Machine Learning Research
Large language models often are called upon to gather news. In this task, researchers found, their ability to find relevant reports is the weakest link.
Business
Web publishers using Cloudflare will soon be able to separately control an AI bot’s access based on what it does, allowing search indexing while blocking AI training or agent activity.
Business
With Llama, Meta marked itself as an open alternative to OpenAI. With its new closed models, Meta now positions itself as a low-cost, high-value competitor.
Machine Learning Research
Moonshot’s latest model leapfrogged the month-old GLM-5.2 and a host of proprietary competitors to finish just behind GPT-5.6 Sol and Claude Fable 5 on many benchmarks.
Letters
A few frontier labs have tried to tell a story of open models being dangerous because they can be used to launch cyberattacks, and of their “safe” proprietary models with strong guardrails being there to defend us.
The Batch Newsletter
The Batch News & Insights: A few frontier labs have tried to tell a story of open models being dangerous because they can be used to launch cyberattacks, and of their “safe” proprietary models with strong guardrails being there to defend us.
Data Points
New rules for Android AI agents in EU. Nemotron 3 Embed takes the embedding model crown. NotebookLM is now Gemini Notebook. How Hugging Face used AI to fight off an AI attack.
Machine Learning Research
Providers of large language models stand to benefit by building models that spur user engagement, but users may bear a cost in undue influence on their world views.
Machine Learning Research
An AI agent proposed new medical uses for established drugs nearly autonomously — uses that were supported by experiments on isolated human cells — with human input only to name diseases to be treated and run the AI-proposed lab experiments.
Tech & Society
A German court ruled that Google can be held liable for defamatory statements generated by the AI Overview that appears at the top of its search results.