CH NEO-ZÜRICH EDITION
WEATHER · OVERCAST 15°C
BLEND OF THE DAY · 07/ROGUE
EST. 2027
THE AEC CYBER MORNING NEWS

PAZ Kaffi

DESIGN · DEMOLITION · CAFFEINE · DISPATCH
EDITION 1002 · 2 October 2026
BROADCAST 04:42 CET
2,400 BROADSHEETS PRINTED
READ TIME · 47 MIN
Gemini 4 Argon: What a $2 Frontier Model Means for a Swiss Office Stack
AI
FRAME · 06:55
02-10-2026

Gemini 4 Argon: What a $2 Frontier Model Means for a Swiss Office Stack

Gemini 4 Argon arrives at $2/$10 per million tokens with a 1M output window. What it means for tender review, nDSG data routing and Swiss office AI stacks.

The most interesting line in Google’s Gemini 4 Argon announcement is not a benchmark. It is a video decoder. According to Google’s own post, Argon agents took an existing Rust port of libgav1, Google’s open-source AV1 decoder, and replaced 32,000 lines of SIMD code. They ran round after round of profile-guided experiments and studied the compiler’s output until safe Rust vectorised on its own. The result is a memory-safe decoder that runs 2.7× faster than the Rust port, with identical video output.

Look at what sits underneath that result. The libgav1 maintainers wrote a reference implementation good enough to measure against. Then someone did the unglamorous work of porting it to Rust in the first place. The model optimised a structure that people had already built carefully. That is the pattern to keep in mind for everything below.

The system behind the signal

Three numbers from the announcement shape any AEC stack decision:

  • Price: an introductory $2 per million input tokens and $10 per million output tokens. Cached input costs 95% less, which works out to $0.10 per million.
  • Output ceiling: up to 1M output tokens in one trajectory, up from 64K.
  • Access: for now, only trusted cyber defenders in Google’s Fairwind Program. Paid API customers and Google AI Ultra subscribers come next.

As CNET points out, today’s reader cannot use Argon yet. CNBC places the launch in the same week that Sundar Pichai signed a voluntary AI safety accord in Washington. Both reports describe a staged rollout, not a switch that has already flipped. Google also says Argon tops long-horizon benchmarks: 77.9% on DeepSWE v1.1, 51.3% on Zapier’s AutomationBench and 91.7% on LVBench for long-video understanding. Those are vendor-reported figures. Treat them as a reason to test, not as a result.

←TODAY: A frontier model launches at $2 / $10 per million tokens with a 1M-token output window, available to cyber defenders first.
→3012: Zurich offices route every dossier through a mapped model graph, and the graph matters more than any single model.
Fulcrum: Cheap tokens only help an office that already knows which documents are allowed to travel, and where they go.

On the desk

Think of a twelve-person studio in Zürich-West in the middle of a Wettbewerb. Its AI stack is rarely designed. It builds up over time: a chat subscription here, an image tool there, a Python script someone wrote in Grasshopper last spring. A model that can digest a full tender dossier and return an equally long answer changes the economics of tender review, code-compliance pre-checks and spec cross-reading against SIA norms. It also adds a dependency. If client drawings go to a US endpoint, the revised Swiss data protection act (nDSG, in force since September 2023) is part of the topology, whether anyone drew it on the diagram or not.

The trade-off, plainly: a 1M-token output window lets a single runaway agent spend up to $10 on one answer, so an office without per-task budgets will find out its real costs at the end of the month.

From where I sit in the late 2070s, the failures that hurt were never the model. They were the quiet dependencies nobody wrote down. Draw your real dependency graph, not the architecture diagram, and find the third single point of failure you didn’t know you had.

Atelier: For an office living with AI, the shift is from “which model is best” to “which document may go to which model, at what cost”. Monday move: make a one-page routing sheet with three columns (data class, permitted model or endpoint, cost cap per task) and pin it next to the PAZ Atelier-Code checklist before anyone opens an API key.

Hack: Price a tender review before you run it, so the token bill is a design input and not a surprise. The figure of 700 tokens per page is a rough estimate for dense spec text, so replace it with a count from your own dossier. The cached line shows why reusing the same dossier across several questions costs a fraction of the first pass.

pages, tok_per_page, out = 400, 700, 20_000   # dossier + clash/summary output
inp = pages * tok_per_page                     # ~280k input tokens
first = inp/1e6*2 + out/1e6*10                 # Argon list price, USD
rerun = inp/1e6*2*0.05 + out/1e6*10            # cached input, 95% off
print(f"first {first:.2f} USD, re-run {rerun:.2f} USD")

The answer is about $0.76 for the first pass and about $0.23 for each re-run. That is cheap enough that the real budget is your reviewer’s time and your data policy, not the tokens. Write the routing sheet before the next tender dossier lands, run the cost line on it, and hold any live test until API access opens.

Source: blog.google

FILED FROM
CO-SIGNERS
PAZ Academy
CONFIDENCE
HIGH
REPRINTS
© PAZ - PARAMETRIC ACADEMY ZURICH · ALL RIGHTS RESERVED

SOURCE · ↗

PAZ Kaffi · multidisciplinary editorial, led by PAZ Academy

⚑ REPORT AN ERROR · SUBMIT A CORRECTION
◂ BACK TO FRONT PAGE · PAZ KAFFI

© 2026 PAZ Academy.