AI News Summary 2026-08-29

AI News Summary 2026-08-29

GAFAM and major AI companies

OpenAI and Google tighten the agent stack

OpenAI’s Codex is now generally available with a new Slack integration, an SDK, and workspace admin tools, while Google introduced flexible usage limits for Gemini Notebook that refresh every five hours and scale with compute-heavy behavior such as longer chats and more sources. Together, the two launches show the same trend from opposite sides of the market: more capable AI assistants, but with more explicit controls for teams and power users. OpenAI Codex GA • Gemini Notebook usage limits

Influencers and technical blogs

Technical digest signals worth following

DeepLearning.AI’s Aug 28 The Batch digest highlights fresh model-watch items such as DeepSeek-V4-Pro Gets Refreshed and GLM-5.3 Makes Cybersecurity Gains, which is a useful signal for readers tracking open and frontier-model movement. Simon Willison’s Aug 27 analysis of Claude Code auto mode is also a strong nearby reference for the current safety debate around coding agents. The Batch tag page • Simon Willison

Generative imagery

Canva and Google keep turning AI images into editable work

Canva’s Magic Layers now appears directly inside Gemini and ChatGPT, letting people turn generated images into fully editable Canva designs. Google Flow also added more creative control, including start and end frames, lower-cost draft renders, and higher-quality exports. The practical result is less “generate a file” and more “ship a usable creative asset.” Canva Magic Layers • Google Flow controls

Chatbots and agents

Claude Code on the web is today’s lead story

Anthropic introduced Claude Code on the web as a browser-based way to delegate coding tasks. It supports parallel sessions, isolated sandboxes, real-time progress tracking, automatic PR creation, and an iOS preview. That makes the product feel less like a terminal tool and more like a browser-native command center for coding work. Claude Code on the web • Claude status

Local AI and serving

Serving stacks keep getting more practical

vLLM’s recent AMD GPU post walks through speculative decoding and related tuning for latency- and throughput-sensitive inference. It’s a good reminder that the local-serving story is still about practical engineering wins: better decode efficiency, more predictable performance, and fewer compromises when you want to run models on your own hardware. vLLM speculative decoding on AMD GPUs