THU, AUGUST 27, 2026
Independent · In‑Depth · Practitioner‑Tested
✎ News

AI News, 27 August 2026

NVIDIA put Groq 3 LPX into full production on 24 August, a dedicated inference accelerator extending Vera Rubin with up to 256 per rack and claiming 4x faster agent responsiveness. OpenAI published Jalapeno benchmarks the next day against the previous generation, and did not test against Vera Rubin. Underneath both sits a design shift: chips optimised for serial agent loops rather than batch throughput.

By AIToolsRecap August 27, 2026 5 min read 107 views
Home Articles News AI News August 27 2026: NVIDIA and OpenAI Both ...
THE 60-SECOND VERSION

● NVIDIA Groq 3 LPX in full production, up to 256 per rack, claiming 4x faster agent responsiveness.

● OpenAI benchmarked Jalapeno against GB300 the day after LPX shipped, and did not test against Vera Rubin.

● The deeper story: chips are being redesigned around agent loops rather than batch throughput.

● 4 days to the 31 August double deadline.

NVIDIA shipped a dedicated inference chip

Groq 3 LPX entered full production on 24 August at Hot Chips. It extends Vera Rubin NVL72 with up to 256 accelerators per rack, splitting the workload — Rubin GPUs handle large-scale context processing while LPX handles latency-critical decode. NVIDIA claims 4x faster responsiveness for agents than the nearest alternative platform, and demonstrated 3,400 tokens per second on Gemma 4 31B at 100K context. Nebius is the first AI cloud to adopt it.

The timing next to OpenAI is the part worth noticing. Jalapeno's benchmarks, published the following day, compared against GB200 and GB300 — the previous generation — and did not test against a Vera Rubin configuration with LPX doing exactly the decode work Jalapeno targets.

Full breakdown: what LPX does, and the claim to be careful with →

And the design shift underneath both

NVIDIA also detailed the Vera CPU: 88 custom Olympus cores across six chiplets, deliberately prioritising single-thread performance for orchestration and tool-calling over raw core count. That is the opposite of a decade of server CPU design, and it follows from what an agent loop actually is — a serial chain where each step waits on the last.

Three designs, one assumption. LPX splits decode from context. Vera optimises orchestration. Jalapeno keeps model state local to cut data movement. All three assume the workload is an agent iterating rather than a batch job completing.

Why this follows from the Linear and Salesforce agent data →

Also this week

  • Thomson Reuters launched its own frontier model for legal work.
  • Mac mini demand has spiked partly on buyers running models locally, with lead times stretching weeks to months. Apple plans to move production to its Houston facility.
  • CoreWeave deployed Spectrum-X Multiplane into production, connecting Vera Rubin racks across parallel switches.

Four days to the double deadline

Date What happens
Aug 31 Claude Sonnet 5 goes $2 to $3 per million input, $10 to $15 output, plus a tokenizer change adding 10 to 35 percent tokens on code
Aug 31 GPT-5.4 and GPT-5.4 mini leave Codex for ChatGPT sign-in users. Still available via API key
Aug 31 kimi-k2.5 and moonshot-v1 sunset, migrate to kimi-k3
Early-mid Sept Grok 4.7 window per Musk. Still no model ID at docs.x.ai
Oct 1 OpenAI vs Apple hearing
Oct 24 deepseek-chat and deepseek-reasoner deprecated
~Nov 21 GPT-5.6 Sol promotional pricing ends, $4/$20 back to $5/$30

Still open

  • Reddit's ChatGPT citations. Down 86 percent in three days after an unannounced retrieval change, with the share moving to vendor documentation. What publishers should check →
  • CVE-2025-62593 in Ray. Anything below 2.52.0 remains exposed. What to check →
  • OpenAI's public S-1. Filed confidentially 8 June, still not on EDGAR.

FAQ

What is NVIDIA Groq 3 LPX?

A dedicated interactive inference accelerator extending Vera Rubin NVL72, in full production since 24 August 2026, with up to 256 accelerators per rack. Rubin GPUs handle context processing; LPX handles latency-critical decode.

Did OpenAI beat NVIDIA?

OpenAI claims Jalapeno delivers 1.5 to 1.9 times more work per watt than GB200 and GB300 — the previous generation. It did not test against Vera Rubin, and both sets of figures come from the companies selling the hardware.

Why are chips being redesigned for agents?

An agent loop is serial: decide, call a tool, wait, read, decide again. Adding parallel cores does not speed up dependent steps, so latency per step matters more than aggregate throughput.

What changes on 31 August?

Claude Sonnet 5 rises to $3 in and $15 out per million with a tokenizer change alongside, and GPT-5.4 and GPT-5.4 mini leave Codex for ChatGPT sign-in users. Both remain available via API.

Tags
AI NewsNVIDIAOpenAIGroqVera RubinAI ChipsAnthropicClaude Sonnet 5Thomson ReutersApple2026

Spot an inaccuracy?

We verify facts before publishing and correct errors promptly. If something in this article is wrong or outdated, let us know.

Report an error →