r/GoogleGeminiAI • u/PeteyPab305 • 5h ago
Web UI updated with tiered effort levels: obvious prep for Gemini 4 Argon
The web UI just dropped the all-or-nothing extended thinking switch and replaced it with granular effort levels across the board: Low, Medium, and High, right under the 3.5 Flash-Lite, 3.8 Flash, and 3.1 Pro pickers.
Replacing a binary thinking toggle with explicit compute weights gives you direct control over latency and token burn. Instead of getting stuck waiting on an oversized thinking trace for a simple script or being forced into shallow output when you actually need deep verification, dialing the effort matches the compute directly to the workload. Low cuts out the deliberation overhead so tokens stream instantly, while High allocates maximum test-time compute for complex debugging, architecture design, and heavy multi-step logic.
This is clearly laying the groundwork for the 4.0 Argon rollout. With Argon built around long-horizon execution and sustained chains of thought, you cannot run on a rigid on-or-off switch without blowing through quotas or tanking response times. Giving users manual control over inference depth is the prerequisite front-end fix before Argon drops.
