Ox Alpha is a multimodal GLM model by Z.ai: Apple Mini and Studio desktops target AI power users

Perplexity moves its computer use agent to on-device models. OpenAI’s Jalapeño may be faster than Nvidia’s best chips. Inside Thomson, Thomson Reuters’ model for law, news, and tax info. T-Rex, a robotics model that knows how to use its hands.

Share
AI robot actively flipping book pages on a desk, showing autonomous interaction with educational content.

We’re fully back from our summer hiatus, and ready for more news!

In today’s edition of Data Points, you’ll learn about our top headlines, and more:

  • Perplexity moves its computer use agent to on-device models
  • OpenAI’s Jalapeño may be faster than Nvidia’s best chips
  • Inside Thomson, Thomson Reuters’ model for law, news, and tax info
  • T-Rex, a robotics model that knows how to use its hands

But first:

Z.ai claims Ox Alpha as its own, will release weights soon

Z.ai confirmed Wednesday that its Ox Alpha model—which launched anonymously over the weekend and surged to the top of OpenRouter's usage charts—is a new iteration of its GLM series, featuring multimodal capabilities including text, image, and video input. The company plans to release the model’s weights soon, allowing developers to build products on top of it, and will keep it free for a week before announcing pricing. The stealth launch, which Z.ai says was inspired by a Chinese film titled “Ox Comes,” follows tactics used by Alibaba and Xiaomi earlier this year: Releasing powerful models without initial credit attribution gives companies freedom to test market response before claiming ownership. Ox Alpha joins a wave of high-performance Chinese AI models released this summer that match or beat US competitors at significantly lower cost, intensifying competition between the two countries. (Bloomberg)

Apple bets big on AI users running local models

Apple updated its Mac Studio with M5 Max and M5 Ultra chips, with the M5 Ultra offering up to 4.3 times the peak AI compute performance of the M3 Ultra, according to Apple. The M5 Max configuration has an 18-core CPU, up to a 40-core GPU, and up to 128GB of unified memory, while the M5 Ultra scales to a 36-core CPU, an 80-core GPU, and up to 512GB of unified memory for running large language models entirely on device. Pre-orders start August 25, with the M5 Max model at $2,499 and the M5 Ultra at $5,499; units ship September 22, and the 512GB memory configuration arrives in late October. Apple also refreshed the Mac mini with two new chips: the M6, starting at $899, and the M5 Pro, starting at $1,699. The devices have yet to be independently benchmarked. The larger memory pools and clustering support for both devices let developers run bigger open-weight models on-device instead of paying for cloud compute. (Apple)

Perplexity updates general-purpose agent to use local models

Perplexity AI released Portable Computer, an AI agent that runs locally on desktops with Nvidia hardware rather than in the cloud. It’s an on-device version of the cloud-based Perplexity Computer agent from February, initially built for Nvidia’s DGX Spark, a machine with a Blackwell-based graphics card, a 20-core CPU, and 128 gigabytes of memory. The agent runs the open-source Qwen 3.8 27B model or a Perplexity-optimized variant called PPLX 27B, uses context compaction to summarize prompts over 100,000 tokens since the model struggles beyond that despite a 256,000-token window, and can escalate hard tasks to a cloud model or connect to services like GitHub with user permission. Perplexity plans to extend support to Windows machines with Nvidia RTX cards and add Nvidia’s Nemotron 3.5 Lightning model. The launch comes as Nvidia reportedly weighs a large investment valuing Perplexity above $30 billion. (Silicon Angle)

Early tests of OpenAI’s custom inference chip show off its speed

OpenAI published initial benchmark results for Jalapeño, its custom AI inference chip, claiming it delivers both higher throughput and lower latency than existing hardware, which typically forces a tradeoff between the two. Tested on GPT-OSS 120B, DeepSeek R1 670B, and Kimi K2.5 1T using the public InferenceX benchmark from SemiAnalysis, Jalapeño delivered 1.5 to 1.9 times more AI work per watt at peak throughput and 1.7 to 3.6 times lower end-to-end latency than comparable systems, with gains up to 4.1 times on highly interactive workloads. The chip is rated at 700 watts but ran at or below 550 watts sustained on tested workloads; OpenAI says internal tests on its own frontier models showed even wider advantages. OpenAI used its own models, including Codex with GPT-Astra, to help design the chip and write optimized kernels, bringing three additional open-weight models to high performance within two months and, for select blocks, generating code 1.5 to 1.8 times faster than human-written implementations. OpenAI plans to begin deploying Jalapeño in its own infrastructure by the end of the year, with a second and third generation already in development, while continuing to use Nvidia and other third-party chips alongside it. (OpenAI)

Thomson Reuters builds specialized model for legal, business, finance, taxes, news, and other knowledge work

Thomson Reuters launched Thomson, its first proprietary large language model, built by fine-tuning an open-weight model with the company’s legal, finance, and news content. The model debuts in Tabular Analysis, a document-review feature inside its CoCounsel Legal assistant, which currently uses third-party frontier models for other tasks. Thomson Reuters says internal tests show the model performs on par with leading models on web-only tasks and slightly better once connected to its own content, though these results haven’t undergone independent validation. The company spent about $40 million over two years on the project but says efficiencies cut the final training run to roughly $450,000. The approach signals a lower-cost path for domain-specific AI: rather than competing head-on with frontier labs, companies with large proprietary datasets can fine-tune open models to outperform general-purpose ones on narrow professional tasks. (Thomson Reuters

Robotics team makes big gains in handheld dexterity

Researchers at Nvidia and UC-Berkeley introduced T-Rex, a robotic manipulation system that lets AI models react to touch in real time rather than just seeing and planning ahead of time. The team built it around a mixture-of-experts transformer architecture that splits control into a slow visual-planning stream and a fast tactile-refinement stream, letting the robot adjust grip force or handle deformable objects on millisecond timescales without slowing down the main policy. To train it, they built a 100-hour dataset pairing 207 household objects with 22 reusable motor primitives, collected on a bimanual robot with tactile sensors on every fingertip, then open-sourced roughly 50 hours of it. Tested across 12 real-world tasks (things like flipping book pages, transferring an egg, or dealing playing cards) T-Rex topped every single task, beating the strongest baseline by more than 30 percent average success rate. Notably, simply bolting tactile sensing onto an existing vision-language-action model made performance worse, not better, suggesting the primary innovation isn’t adding touch data but architecting how and when a robot listens to it.(GitHub)


Want to know more about what matters in AI right now? 

Read the latest issue of The Batch for in-depth analysis of news and research.

Last week, Andrew talked about the essential skills for building and deploying AI applications, focusing on LLM foundations, grounding models with data, and evaluation-driven development.

“In my experience, the most important trait that distinguishes someone great at building AI systems is whether you can drive a disciplined evals/error analysis loop to drive development. This allows you to repeatedly focus your effort on directions that are more likely to be fruitful. I’ve found this to be a tricky skill to master, because the right approach varies significantly by project and even according to the stage of the project.”

Read Andrew’s letter here.

Other top AI news and research covered in depth:


A special offer for our community

DeepLearning.AI’s first-ever subscription plan for our entire course catalog includes foundational classics like the Machine Learning and Deep Learning specializations plus seminars on the latest tools and frameworks you need.

As a Pro Member, you’ll immediately enjoy access to:

  • Nearly 200 short and long AI courses from Andrew Ng and industry experts
  • Labs and quizzes to test your knowledge
  • Projects to share with employers
  • Certificates to testify to your new skills
  • A community to help you advance at the speed of AI

Enroll now to lock in a year of full access for $25 per month paid upfront, or opt for month-to-month payments at just $30 per month. Both payment options begin with a one-week free trial.

Explore Pro’s benefits and start building today!

Try Pro Now!