AI News Summary 2026-07-15

.. lang: en

AI News Summary July 15, 2026

Today’s overarching theme was clear: companies are treating AI as an operational layer with metrics, governance, and real-world deployment; at the same time, agents are moving beyond chat and approaching the operating system, and local serving continues to be refined for low-latency workloads.

GAFAM and Major AI Companies

How to Manage AI Investments in the Agentic Era

OpenAI published a guide for investing with greater confidence in the agentic era and states that the cost per million tokens fell by 97% from GPT-4 to GPT-5.4; furthermore, GPT-5.6 improves the Artificial Analysis Coding Agent Index with 54% fewer output tokens and 57% less time per task. OpenAI article · OpenAI news

How Gemini is speaking the language of Southeast Asia

Google says that active Gemini users in Southeast Asia have more than doubled in the last year and that Gemini Spark is being rolled out in local languages for Gemini Advanced subscribers. Google article · Innovation & AI

Influencers and Tech Blogs

TIL: Using uvx in GitHub Actions in a cache-friendly way

Simon Willison shares a practical pattern for caching uvx in GitHub Actions using UV_EXCLUDE_NEWER, with the goal of avoiding repeated downloads from PyPI on every run. post · Atom feed

Generative imaging

Introducing Muse Image and Muse Video

Meta introduced Muse Image for following instructions, editing, and composition with multiple references, and Muse Video with native audio support. Although it isn’t the freshest news of the day, it remains the most solid highlight to emerge from the generative image source review. Meta AI · blog

Chatbots and Agents

Personal Computer Is Here

Perplexity has begun rolling out Personal Computer, an expansion of Perplexity Computer that brings multimodal orchestration to local files, native apps, connectors, and the web, while keeping the user in control of sensitive actions. Perplexity blog · hub

On-Premises AI and Serving

vLLM x TileRT: Specialized Decode for Latency-Critical Serving

vLLM introduces a pluggable decode path with TileRT for low-latency workloads, while keeping the prefill and the rest of the vLLM stack intact. vLLM article · blog