LLMs Take Out the Agent’s Trash: Researchers at Xiaohongshu detail LLM-based memory management technique for agents
AI agents typically compact the contents of their context windows by summarizing or deleting the oldest material.
AI agents typically compact the contents of their context windows by summarizing or deleting the oldest material.
DeepSeek’s flagship model graduated from preview with improved performance plus the harness the model was benchmarked in.
Developers who already track models’ cost and accuracy have good new reasons to pay closer attention to a third essential factor: speed.
Z.ai’s latest flagship model effectively ties open-weights leader Kimi K3 on Artificial Analysis’ index of intelligence benchmarks.
How have software engineering fundamentals changed with agentic coding?
The Batch News & Insights: How have software engineering fundamentals changed with agentic coding?
Perplexity moves its computer use agent to on-device models. OpenAI’s Jalapeño may be faster than Nvidia’s best chips. Inside Thomson, Thomson Reuters’ model for law, news, and tax info. T-Rex, a robotics model that knows how to use its hands.
GLM-5.3 jumps to top of open weight leaderboards. Harvey’s Tenet offers law firms a specialized model. OpenAI introduces safety program for data-conscious customers. New Google research gives the transformer a makeover.
Most speech-to-text systems transcribe speech in a single pass, which makes them unable to correct errors in their outputs. Researchers built a system that allows for interactive corrections.
Open models are getting larger and more capable. A few weeks ago, Moonshot AI announced Kimi K3, the largest and best-performing open weights model yet. Last week, Alibaba answered by releasing weights for a giant of its own.
Anthropic introduced invisible, machine-readable signals that text and images were generated by Claude.
Once a lab that produced mid-tier models, SpaceXAI has steadily improved. It just built one of the most capable models in the world while keeping prices relatively low.